🏡


  1. August 20, 2026
    1. 🔗 HexRaysSA/plugin-repository commits sync repo: +1 plugin, +3 releases, ~12 changed rss
      sync repo: +1 plugin, +3 releases, ~12 changed
      
      ## New plugins
      - [ida-nexus](https://github.com/hexrayssa/ida-nexus) (0.6.2)
      
      ## New releases
      - [AutoRE](https://github.com/a1ext/auto_re): 2.3.0
      - [efiXplorer](https://github.com/rehints/efixplorer): 6.3.0
      
      ## Changes
      - [ida-codemode](https://github.com/hexrayssa/ida-codemode):
        - 0.6.1: download URL changed
        - 0.6.0: download URL changed
        - 0.5.3: download URL changed
        - 0.5.2: download URL changed
        - 0.5.1: download URL changed
        - 0.5.0: download URL changed
        - 0.4.1: download URL changed
        - 0.4.0: download URL changed
        - 0.3.2: download URL changed
        - 0.3.1: download URL changed
        - 0.3.0: download URL changed
        - 0.2.0: download URL changed
      
    2. 🔗 anthropics/claude-code v2.1.238 release

      What's changed

      • Added a keybindingFlavor setting: set it to "readline" to make Ctrl+W in the prompt delete back to the previous whitespace, as in Bash; the default ("classic") is unchanged
      • Plugin marketplaces: headersHelper on a url marketplace or a catalog entry runs a command that mints HTTP headers (e.g. a short-lived token) for catalog and same-origin archive fetches
      • A catalog entry's headersHelper runs only when you install or update that plugin, after its command is shown; claude plugin install/update ask [y/N] (or pass -y)
      • Added claude self-hosted-runner --defer-shutdown-max-min <minutes>: on SIGTERM, keep serving attached sessions, park what is left after that many minutes, then exit
      • Added claude self-hosted-runner --proxy-authorization-command / --proxy-authorization-file for egress proxies that require a freshly issued Proxy-Authorization header on every connection
      • Fixed unbounded memory growth in long interactive sessions: subagent tool results are now released once they leave the recent display window
      • Fixed custom, project, and plugin output styles drifting back to the default voice mid-session
      • Fixed CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION=true not keeping prompt suggestions on when your account is near, but not over, its usage limit
      • Fixed worktree-isolation Bash refusals telling you to remove a redirect when the command had none
      • Fixed self-hosted runners occasionally being removed by the server after a single slow or lost poll request, handing their healthy session to another runner
      • Fixed MCP elicitation dialogs showing nothing for URLs longer than 4,096 characters, and permission prompts dropping the "don't ask again" option when the project path didn't fit the terminal width
      • Fixed leftover /tmp/claude-*-cwd files when a Bash command is killed, times out, or is interrupted
      • Fixed held Backspace being ignored on terminals that send Ctrl+H for Backspace when keystrokes arrive in large bursts (slow SSH/mosh links)
      • Fixed text-wrapping in permission prompt diffs: lines containing wide multi-code-point characters (such as emoji) or tabs are no longer clipped
      • Fixed killing a suspended (Ctrl+Z) session sometimes leaving the terminal in bracketed-paste mode with the cursor hidden
      • Fixed stdio MCP servers receiving a server/discover request before initialize, forcing lazy servers to start their backend on every session open
      • Fixed a proxy's refusal of a connection being reported as a generic network error instead of naming the proxy
      • Fixed the /model and /effort cache-miss warning appearing when the prompt cache had already expired
      • Fixed per-task Stop from the Remote Control tasks panel doing nothing on CLI-hosted sessions
      • Fixed remote sessions exiting when a client delivered a user message without a valid role
      • Fixed Remote Control sessions started by claude remote-control inheriting session-scoped environment variables from the launching shell
      • Fixed a Remote Control session whose process crashed staying unavailable until claude remote-control was restarted; it can now be reused when you next message it
      • Fixed Remote Control messages sent from the web or Desktop while Claude is mid-turn disappearing from the transcript after the turn finishes
      • Fixed Remote Control model picks made on a phone or web not updating the model shown in the terminal
      • Fixed Remote Control disconnecting with "login expired" when a brief network hiccup delays renewing your sign-in; it now retries and stays connected
      • Fixed Remote Control reporting a failed reconnect on sign-out; signing out now ends the session with a clear message
      • Fixed ListAgents/SendMessage reporting "Remote Control is not connected" in sessions run by claude remote-control (server mode) or Desktop/IDE hosts; they now list and reach Remote Control peers
      • Fixed ListAgents and SendMessage exposing the idle worker that the agent view pre-warms for your next background session; it now appears only once a task claims it
      • Cross-session messaging: sending to a session on this machine that refuses inbound messages (e.g. crossSessionInbound: "refuse") now reports "refused" to the sender instead of a silent success
      • Cross-session messaging: a session whose inbox drops your messages (rate limit or full queue) now tells your session, instead of the messages vanishing silently
      • Improved startup: bare claude starts sooner on macOS
      • Improved Bash tool permission checking for zsh-specific syntax in shell conditionals
      • Improved Remote Control connection resilience: brief HTTP 403 refusals from a network edge, VPN, or proxy are now tolerated for up to 3 minutes, with the refusing party named when a block persists
      • Improved startup responsiveness: the automatic update check now runs about 10 seconds after launch instead of competing with startup for CPU
      • Updated the bundled claude-api skill for the Managed Agents Aug 19 release: web search/fetch domain settings and memory stores on self-hosted sandboxes
      • Changed Ctrl+L and Cmd+K in fullscreen to always just repaint — the double-press /clear shortcut was removed, and 1-row nvim terminals no longer trigger automatic /clear loops
      • Changed claude mcp list and claude mcp get to show disabled servers as ⊘ Disabled instead of connecting to them for a health check
      • MCP headersHelper in a project .mcp.json, and inline MCP servers in project or --add-dir agent files, now require that folder's trust dialog to have been accepted (also under claude -p)
      • MCP headersHelper from a project .mcp.json, plugin, or agent file runs without inherited credential env vars; user, managed and claude.ai-scope helpers now run from the Claude config dir
    3. 🔗 PrimeIntellect-ai/prime-agent Beta (v0.7.4-beta.531.1.c75a637) release

      Automated beta build from main (c75a637b00d3b52762841e72efe92289a0d55b49).

    4. 🔗 The Pragmatic Engineer The Pulse: Meta’s self-inflicted resignation-wave rss

      Hi, this is Gergely with a bonus, free issue of the Pragmatic Engineer Newsletter. In every issue, I cover Big Tech and startups through the lens of senior engineers and engineering leaders. Today, we cover one out of four topics from last week 's The Pulse issue. Full subscribers received the article below seven days ago. If you 've been forwarded this email, you can subscribe here .

      Two months ago, I covered how Meta seemingly deliberately started destroying its once-standout engineering organization. The company did 10% layoffs at a time when revenue and profits hit an all-time high, and reassigned about 20-30% of software engineers to data labeling with basically no notice.

      The result was that a good chunk of software engineers at Meta - even those not reassigned - were starting to interview elsewhere:

      altWhy so many engineers at Meta are looking for a way out, visualized. Source:Why is Meta destroying its engineering organization?

      I said - and still maintain - that the layoffs were an unforced, self- inflicted error on Meta's part. Coupled with forced reassignment, they struck disaster. And I've now gathered new details that confirm that Meta is bleeding top engineering and product management talent:

      Some layoffs reversed, in the 11th hour

      In the UK, mass layoffs require a notification to those potentially affected, before the cuts can happen. A few weeks after this notification, a good chunk of folks were "un-notified," I confirmed. But those on notice had already started to look for jobs, obviously.

      Retainer equity grants offered to IC6+ engineers resigning

      Meta started to offer large retainer equity grants to those resigning and leaving for Google, Anthropic, and OpenAI: a practice not done before. Talking with long-time Meta engineers, Meta simply didn't make counteroffers or negotiate when an engineer handed in their resignation. But this practice has been abandoned. Here's a story of a senior, long-tenured engineer I confirmed to have happened:

      • This dev was force-reassigned to one of the AI data labeling teams
      • They thought "oh no, this is not what I want to do" and started to interview
      • Got a Google offer, but compensation was below their current one. Google basically lowballed them.
      • Decided to leave anyway. Told their new manager they were resigning.
      • New manager asked for some time, then presented a large, one-off, retainer equity grant, vesting over 3 years
      • The engineer went back to Google, telling the company, unfortunately, they are a key engineer and showed them the retainer amount Meta is paying for them to stay
      • Google upped their offer, out-bidding Meta's retainer offer
      • The engineer who was desperate to leave - even for lower compensation! - was thrilled, and happily left for Google.

      Retainer equity seems to be offered for most IC6 engineers who resigned in the last few weeks. I talked with seven such people, who were all IC6 (staff-level engineer) and IC7 (principal-level engineers), who shared more details. They are not aware of IC4 (mid-level engineer) or IC5 (senior-level engineers) getting such offers.

      Retainer equity offered is $400K to $1M+ , and vests over 3 years, for the folks who got these offers. The $1M+ offers were all for engineers with Anthropic or OpenAI offers, and I confirmed a $400K and a $600K grant for engineers with an offer for other, smaller AI startups.

      A current Meta engineer who is interviewing asked me if these folks needed to present their offer letter to get a counteroffer. I asked two of these folks with counteroffers, and neither had to do so, but they both indicated they were handing in their resignation. Upon receiving this notification, their director, together with HR, offered a discretionary retainer equity, should they stay.

      Anthropic and OpenAI on a hiring spree from Meta

      Anthropic and OpenAI seem to be on a hiring spree, signing Meta engineers that Meta wants to retain now that it is too late! I have confirmed that three engineers who received an offer from Anthropic were also offered large, $1M+ equity grants (vesting over 4 years) as counter-offers to stay. Two out of the three rejected it and joined Anthropic, and the third one initially accepted the offer and then still left for Anthropic a month later, forfeiting this grant.

      OpenAI is having similar success, and the two AI labs seem to be the main destination for now ex-Meta engineers to bounce to. It makes sense: these companies can match Meta's total compensation, and although neither OpenAI nor Anthropic is publicly traded, both companies organize secondary equity sales, and their respective IPOs are likely to happen in 3-12 months' time.

      AI startups also successful in hiring infra experts from Meta

      I had an in-depth discussion with another long-tenured AI infra engineer at Meta, who was considering whether to stay, or to go. This person wrote up their situation, and shared it with me, writing:

      "With the current situation at Meta (layoffs, forced drafting to do AI work, general sense of disruption at work, low moarale, and almost zero productivity across teams, plus multiple reorgs) I started to look around, on the market, a few weeks before the 20 May layoffs. I figured I have a 50% chance of being let go.

      I did not get laid off, but I got offers from startups, and one from a high- growth AI startup I like, and also an offer from Google.

      Now, the pros to staying a Meta:Context, relationshipsScale that few others haveBig Tech perks, relatively "chill" WLB in terms of oncallHigh TC

      The cons:Terrible morale, hard to "move" work through mostly unmotivated teams all around meRe-orgs continue, I have very low confidence in middle management (Directors, VPs) to navigate this churnFurther layoffs are likely: it's clear that there's more bloat, surely Zuck will do another round in 2027. And no one is safe during a layoffA 'doom loop' for anyone who stays: low morale -> unmoitivated bith highly compensated engineers -> a bloated organization that moves slow"

      After a new reorg, this engineer got a supportive manager, who made it clear they would help them thrive inside of Meta. So the engineer rejected both the AI startup's offer and that of Google, deciding to stay… but only for a month!

      The AI startup's founder spent more time with this engineer, increased the equity in the package significantly, and convinced this engineer that at the startup they matter as a person, while at Meta they would be doing soulless work, worrying about the next layoff, probably early 2027.

      So this engineer also handed in their resignation. And it all started with the layoffs: the engineer assumed there was a 50% chance of being let go and wanted more career stability!

      Those deciding to not resign: stay for the money?

      There are engineers staying - at least until the end of the year, to wait for the stock refershers -- which we'll talk about in just a moment.

      Another engineer I talked to decided to quit Meta with morale so low, and start his own thing. Management convinced him to stay until the end of the year, hinting at large equity refreshers being handed out around January, which could be in the $1M+ range for this specific engineer.

      The 2022 stock price causing a TC decrease for many?

      One reason there's speculation about large equity refreshers being handed out at the end of the year is that Meta's stock price was very low in 2022 and 2023, when equity refreshers and new hire grants were handed out. In both March 2022 and March 2023, the stock price was around $200 - which is about a third of this year's stock price (which has been fluctuating between $520- 680, currently sitting at $540.) By March 2024, the stock price rose to $500, and then to $580 in March 2025.

      Because of this, those who joined Meta in 2022 have seen their equity grants nearly triple, and the 2023 equity is worth nearly 3x as well. But for these lucky folks, total compensation is set to drop by the end of this year, and their equity refresher would need to be higher than in past years to avoid a compensation drop.

      Inside of Meta, engineers who I talked to are split between expecting further attrition (those leaving whose total compensation drops steeply) and those who think that there will be more attrition unless there are higher-than-usual equity refreshers.

      A "mercenary" culture at Meta?

      I wonder if Meta made a massive mistake by treating engineers as "commodities" and turning their culture into a "mercenary" one. Between March and June, Meta treated engineers as replaceable commodities that could be thrown out or moved between teams, while stripping them of any autonomy. The reassignments to AI labeling were not explained, nor were managers or individuals asked about preferences: from above, someone said that Alice and Bob , starting on Monday, are no longer with the team, but are training up Meta AI. Never mind that these were the two most experienced engineers on the team, or that Bob was the team's infra expert, and Alice the "fixer" engineer on the team.

      And all this damage was done, for what? It was to allow Meta to restart its AI coding model development efforts, and release Meta Muse Spark: an AI model that is currently the 7th most capable AI model as per Artificial Intelligence Analysis, tied with Grok 4.5. Credit where it's due: with this model, Meta is now well ahead of Google's Gemini in capability.

      With the damage done, there is little motivation for anyone to stay at Meta - unless they are on a visa, or if it's about the money. You can have an OK engineering culture with a team who is mostly interested in making more money than what they would elsewhere: but this is more of a mercenary culture, and companies run by mercenaries are more easily out-executed by teams where people believe in the mission of the company.

      If you work at a company that would love to hire from Meta: now is (still) your opening! So go for it.

      And if you're at Meta: you are so in-demand, outside of the company, at places where engineering is still treated as a profit center. As Ryan Nystrom, currently building AI at Notion, put it:

      "Meta friends: the industry is exhilarating right now. Get out. Go build something. The handcuffs are an illusion.

      I'm having the time of my life at Notion. Come, and have fun again."

      Read the full issue of last week 's The Pulse, or check out this week 's The Pulse. This week's issue covers:

      1. More on the "great engineering leader career break." The industry is changing fast, and the VPE and CTO roles also need to adapt. And don't forget that these are the roles from which you can drive change that reorganizes engineering in ways that work better.
      2. We need to talk about migrations with AI. Asana needed to migrate off testing framework Enzyme, but it meant doing a massive rewrite of test cases. With AI, the project was completed in two weeks: without AI, this work would surely have been kicked down the road. Airbnb and Uber share similar stories, and AI seems like a superb fit for framework migrations.
      3. Are AI startups making the Gartner Magic Quadrant irrelevant? Gartner ranked AWS, Microsoft and IBM above Anthropic, Cursor and OpenAI in their "AI code modernization tools" ranking. This is most likely because the first three pay large sums of money to Gartner, but AI labs and vendors refuse to pay this "Gartner tax."
      4. Industry Pulse. Another hours-long GitHub outage, GitHub alternatives are here and fighting for market share, Slack launches Slack Code, text generated by Claude to be watermarked, and Uber open sources SubmitQueue.

      Read the full The Pulse.

    5. 🔗 HexRaysSA/ida-nexus v0.6.2 release

      Full Changelog : v0.6.1...v0.6.2

    6. 🔗 Anton Zhiyanov Going freestanding rss

      Creating a subset of Go that translates to C (which I named Solod) was never my end goal. I liked writing C code with Go, but without the standard library it felt pretty limited. So the next logical step was to port Go's stdlib.

      At some point I decided to make as many packages as possible freestanding — independent of any libc implementation or specific OS runtime. That went pretty well. Solod now has 37 standard library packages, and 31 of them work in freestanding mode.

      This post describes the techniques I used to get there. There's nothing genuinely novel, and if you're experienced with C, you probably already know all of them. Still, I think it's useful to document the approach — both for me and for anyone interested.

      Freestanding modeHeadersBuiltinsMemoryAtomicsPure CAllocationValuesHooksHosted- onlyTestingFinal thoughts

      Freestanding mode

      C has two types of environments. In a hosted environment, you get the full standard library — either the one required by the C standard or, even better, POSIX. In a freestanding environment, you get almost nothing.

      The compiler tells you which one you're in:

      #if __STDC_HOSTED__
      // libc is available
      #else
      // you're on your own
      #endif
      

      Pass -ffreestanding and link with -nostdlib, and that's it: you no longer have printf, or malloc, or even memcpy. There is no entropy source, no file system operations, and no clock. If libc itself is "hard mode", this is "impossible".

      Despite its limitations, freestanding mode can be really useful for microcontrollers, WebAssembly sandboxes, kernels, and anything else without an operating system to rely on.

      Freestanding headers

      Freestanding does not mean "just the C language". The C standard guarantees some headers even without libc, because they define types and macros rather than functions:

      float.h   stdalign.h  stdbool.h  stdint.h
      limits.h  stdarg.h    stddef.h   ...
      

      Everything that requires actual function implementations is gone:

      assert.h  math.h   stdlib.h  time.h
      errno.h   stdio.h  string.h  ...
      

      To reflect the hosted/freestanding split, let's introduce builtin.h, a common header included in every standard library package:

      #if __STDC_HOSTED__
      
      #include <assert.h>
      #include <inttypes.h>
      #include <stdalign.h>
      #include <stdbool.h>
      #include <stdint.h>
      #include <stdio.h>
      #include <stdlib.h>
      #include <string.h>
      
      #define so_build_hosted
      
      #else
      
      #include <stdbool.h>
      #include <stdint.h>
      #include <stdalign.h>
      #include <stddef.h>
      
      #endif  // __STDC_HOSTED__
      

      Individual packages follow the same approach: branch on so_build_hosted to distinguish between the hosted and freestanding implementations.

      Compiler builtins

      GCC and Clang implement some C standard functions without relying on libc. These are known as compiler builtins.

      __builtin_trap causes the program to terminate abnormally. You can use it to implement poor man's assertion and panic:

      #ifdef so_build_hosted
      
      #define so_panic(msg)                                     \
          do {                                                  \
              fprintf(stderr, "panic: %s\n  %s:%d (func %s)\n", \
                      msg, __FILE__, __LINE__, __func__);       \
              exit(1);                                          \
          } while (0)
      
      #else
      
      #define assert(cond)                   \
          do {                               \
              if (!(cond)) __builtin_trap(); \
          } while (0)
      
      #define so_panic(msg)     \
          do {                  \
              (void)msg;        \
              __builtin_trap(); \
          } while (0)
      
      #endif // so_build_hosted
      

      From now on, I'll mainly show the freestanding versions and omit the hosted versions to keep things simple.

      The __builtin_alloca function allocates memory on the stack. Its bounded wrapper limits the size of each allocation:

      #define alloca __builtin_alloca
      
      // The maximum size that can be allocated
      // with alloca (64 KB by default).
      #ifndef SO_MAX_ALLOCA_SIZE
      #define SO_MAX_ALLOCA_SIZE (64 << 10)  // in bytes
      #endif
      
      #define so_alloca(size) ({                                \
          size_t _size = (size_t)(size);                        \
          if (_size > SO_MAX_ALLOCA_SIZE)                       \
              so_panic("alloca: size exceeds maximum allowed"); \
          _size ? alloca(_size) : NULL;                         \
      })
      

      Memory operations

      The memxxx functions from string.h have matching builtins too, so you might expect a freestanding build to provide them for you:

      // int memcmp(const void* lhs, const void* rhs, size_t n);
      #define memcmp __builtin_memcmp
      
      // void* memcpy(void* dst, const void* src, size_t n);
      #define memcpy __builtin_memcpy
      
      // void* memmove(void* dst, const void* src, size_t n);
      #define memmove __builtin_memmove
      
      // void* memset(void* dst, int ch, size_t n);
      #define memset __builtin_memset
      

      Unfortunately, there's no free lunch here.

      __builtin_memcpy is not a separate memcpy implementation. If n is small and known at compile time, the compiler expands it into a few loads and stores. But if n is large or only known at runtime, it emits a call to the real memcpy — the same symbol that libc would provide.

      Even worse, you don't need to mention memcpy explicitly to use it. Suppose you copy a large struct like this:

      typedef struct { char buf[4096]; } Big;
      
      void copy(Big* a, const Big* b) {
          *a = *b;
      }
      

      When you compile the code for aarch64-freestanding, the object file contains an undefined reference to memcpy. Zero-initializing a local array produces the same issue with memset. Neither name appears in the source; both are introduced by the compiler.

      So the freestanding environment must still provide memcpy, memmove, memset, and memcmp for memory operations to work in the general case.

      WebAssembly covers three of the four: memcpy, memmove, and memset lower to the memory.copy and memory.fill instructions. There is no instruction for comparison, so memcmp stays a real function call even there. On other targets, the toolchain often provides all four, as zig cc does (even with -nostdlib). If it doesn't, provide a plain C implementation:

      #undef memcpy
      void* memcpy(void* dst, const void* src, size_t n) {
          unsigned char* d = dst;
          const unsigned char* s = src;
          while (n--) *d++ = *s++;
          return dst;
      }
      
      #undef memset
      void* memset(void* dst, int ch, size_t n) {
          unsigned char* d = dst;
          while (n--) *d++ = (unsigned char)ch;
          return dst;
      }
      
      #undef memmove
      void* memmove(void* dst, const void* src, size_t n) {
          // omitted for brevity
      }
      
      #undef memcmp
      int memcmp(const void* lhs, const void* rhs, size_t n) {
          const unsigned char* l = lhs;
          const unsigned char* r = rhs;
          for (; n--; l++, r++) {
              if (*l != *r) return *l - *r;
          }
          return 0;
      }
      

      The defines (#define memcpy __builtin_memcpy and others above) are still worth keeping, even if you implement the functions yourself. This way, the compiler can still use its own implementation when applicable.

      Fun fact: at -O2 and above, GCC can fold your custom memcpy implementation back into a call to memcpy, which is infinite recursion. -ffreestanding prevents this because it implies -fno-builtin, but the guarantee is weak. You can use -fno-tree-loop-distribute-patterns to disable this behavior for good.

      The rest of string.h is not covered. No compiler provides memchr or strlen, so those you always have to write yourself — more on that below.

      Atomic operations

      Another useful group of compiler builtins is __atomic_xxx, which provide atomic, thread-safe memory access. They operate on regular objects rather than _Atomic objects:

      // so_atomic_load atomically loads the value at p.
      #define so_atomic_load(p) \
          (__atomic_load_n((p), __ATOMIC_SEQ_CST))
      
      // so_atomic_store atomically stores v at p.
      #define so_atomic_store(p, v) \
          (__atomic_store_n((p), (v), __ATOMIC_SEQ_CST))
      

      This makes porting Go's sync/atomic types straightforward. All types — atomic integers, unsigned integers, booleans, and pointers — use the same two load/store macros:

      // Bool is an atomic boolean value. The zero value is false.
      typedef struct atomic_Bool {
          bool v;
      } atomic_Bool;
      
      // Load atomically loads and returns the value stored in x.
      bool atomic_Bool_Load(atomic_Bool* x) {
          return so_atomic_load(&x->v);
      }
      
      // Store atomically stores val into x.
      void atomic_Bool_Store(atomic_Bool* x, bool val) {
          so_atomic_store(&x->v, val);
      }
      

      A separate atomic_Bool type isn't strictly required — atomic_Bool_Load and atomic_Bool_Store would work with a plain bool*. Still, it can be useful. With a plain pointer, *x = true creates a silent data race that looks like ordinary code, while the wrapper makes you explicitly write x->v = true.

      No stdatomic.h include is needed. However, the CPU must natively support the integer width you use. For example, a 64-bit atomic on a 32-bit target becomes a call to libatomic instead of a single instruction. But that's a different story.

      Pure C implementations

      If the compiler doesn't provide an implementation, you have to write one yourself. Preferably, use the libc name so all call sites remain unchanged.

      A good example is memchr, which is required by bytes.IndexByte. There is a __builtin_memchr, but it is not an implementation, so you need to provide your own:

      #ifndef so_build_hosted
      // memchr implementation for freestanding environments.
      static inline void* memchr(const void* s, int c, size_t n) {
          const unsigned char* p = s;
          unsigned char target = (unsigned char)c;
          while (n--) {
              if (*p == target) return (void*)p;
              p++;
          }
          return NULL;
      }
      #endif
      

      Some of these DIY implementations aren't trivial, of course. Fortunately, Go's standard library includes many standalone algorithms, such as the string-to- number conversion functions in strconv or integer math in math/bits. Porting them to C is almost mechanical:

      // Go version.
      const m3 = 0x00ff00ff00ff00ff
      
      // ReverseBytes32 returns the value of x
      // with its bytes in reversed order.
      func ReverseBytes32(x uint32) uint32 {
          const m = 1<<32 - 1
          x = x>>8&(m3&m) | x&(m3&m)<<8
          return x>>16 | x<<16
      }
      
      
      
      // C version.
      static const int64_t m3 = 0x00ff00ff00ff00ff;
      
      uint32_t bits_ReverseBytes32(uint32_t x) {
          const int64_t m = ((int64_t)1 << 32) - 1;
          x = ((x >> 8) & (m3 & m)) | ((x & (m3 & m)) << 8);
          return (x >> 16) | (x << 16);
      }
      

      Memory allocation

      Memory allocation calls for a different technique. A naive approach would be to implement a freestanding malloc that uses a static buffer:

      extern char so_heap[SO_HEAP_SIZE];
      extern size_t so_heap_offset;
      
      static inline void* malloc(size_t size) {
          // Simplified version without alignment.
          if (size > SO_HEAP_SIZE - so_heap_offset) {
              return NULL;
          }
          void* ptr = &so_heap[so_heap_offset];
          so_heap_offset += size;
          return ptr;
      }
      

      It might be sufficient for testing, but I'd avoid using it in production.

      Instead of reimplementing malloc, let's remove the need for it, and make the caller bring the memory. Start with an allocator interface, so callers don't depend on a specific implementation:

      // Allocator defines the interface for memory allocators.
      // Simplified version without Realloc and alignment.
      typedef struct {
          void* self;
          so_R_ptr_err (*Alloc)(void* self, so_int size);
          void (*Free)(void* self, void* ptr, so_int size);
      } mem_Allocator;
      

      What's with the so-types?

      so_int is an integer of the target width:

      #if SIZE_MAX == 0xFFFFFFFFu
      typedef int32_t so_int;
      #else
      typedef int64_t so_int;
      #endif
      

      so_String is a pointer to the underlying string bytes and their count:

      typedef struct {
          const char* ptr;
          so_int len;
      } so_String;
      

      so_Error is an interface value that wraps the error data:

      typedef struct {
          void* self;
          so_String (*Error)(void* self);
      } so_Error;
      

      so_R_ptr_err is a result-type implementation for a (pointer + error) pair:

      typedef struct {
          void* val;
          so_Error err;
      } so_R_ptr_err;
      

      There are other similar types like so_R_int_err (int + error) or so_R_f32_bool (float32 + bool).

      Then provide an arena allocator, which is freestanding by design:

      // Arena is a memory allocator that bump-allocates
      // linearly within a fixed buffer.
      typedef struct {
          so_Slice buf;
          so_int offset;
      } mem_Arena;
      
      mem_Arena mem_NewArena(so_Slice buf) {
          return (mem_Arena){.buf = buf};
      }
      
      so_R_ptr_err mem_Arena_Alloc(void* self, so_int size) {
          // Simplified version without alignment.
          mem_Arena* a = self;
          assert(size > 0 && "mem: invalid allocation size");
          if (size > so_len(a->buf) - a->offset) {
              return (so_R_ptr_err){.val = NULL, .err = mem_ErrOutOfMemory};
          }
          void* ptr = &so_at(so_byte, a->buf, a->offset);
          a->offset += size;
          return (so_R_ptr_err){.val = ptr, .err = (so_Error){}};
      }
      
      void mem_Arena_Free(void* self, void* ptr, so_int size) {
          // Free in arena is a no-op.
          (void)self; (void)ptr; (void)size;
      }
      
      void mem_Arena_Reset(void* self) {
          mem_Arena* a = self;
          a->offset = 0;
      }
      

      Usage example:

      typedef struct Point {
          so_int x;
          so_int y;
      } Point;
      
      // Prepare the arena.
      so_byte data[1024];
      so_Slice buf = {.ptr = data, .len = sizeof(data)};
      mem_Arena arena = mem_NewArena(buf);
      mem_Allocator alloc = {
          .self = &arena,
          .Alloc = mem_Arena_Alloc,
          .Free = mem_Arena_Free};
      
      // Allocate a Point. mem_Alloc is a macro that calls
      // the Alloc "method" and panics on failure.
      Point* p = mem_Alloc(Point, alloc);
      p->x = 11;
      p->y = 22;
      

      On a freestanding target, an arena is a better choice than a buffer-backed malloc, because the caller decides how much memory is available and when it's released.

      Values, not pointers

      Constructor functions in Go typically return a pointer:

      // A string reader.
      type Reader struct {
          s        string
          i        int64 // current reading index
          prevRune int   // index of previous rune; or < 0
      }
      
      // NewReader returns a new Reader reading from s.
      func NewReader(s string) *Reader {
          return &Reader{s, 0, -1}
      }
      

      This roughly translates to the following code, using the memory allocator from the previous section:

      // A string reader.
      typedef struct {
          so_String s;
          int64_t i;
          so_int prevRune;
      } strings_Reader;
      
      // NewReader returns a new Reader reading from s.
      // The returned reader is allocated; the caller owns it.
      strings_Reader* strings_NewReader(mem_Allocator alloc, so_String s) {
          strings_Reader* r = mem_Alloc(strings_Reader, alloc);
          r->s = s;
          r->i = 0;
          r->prevRune = -1;
          return r;
      }
      

      Rather than blindly following Go idioms, it's better to get rid of allocations altogether and return a value:

      // NewReader returns a new Reader reading from s.
      strings_Reader strings_NewReader(so_String s) {
          return (strings_Reader){.s = s, .prevRune = -1};
      }
      

      This isn't a technique specific to writing freestanding code, but rather a useful practice for pretty much any C library.

      Target hooks

      Some things you can't write in a target-agnostic way at all. Only the target knows how to print a byte, read the clock, or generate a random number; these all depend on the hardware.

      What you can do is declare functions (hooks) and let the user's code define them:

      Hook | Description
      ---|---
      so_write_out | send some bytes to the output
      so_crand_read | read some random bytes
      so_time_wall | get the current wall clock time
      so_time_mono | get the current monotonic time
      so_time_sleep | pause for a given duration

      Then the user can call specific APIs available on their hardware:

      so_int so_write_out(const uint8_t* buf, so_int size) {
          return board_uart_write(buf, size);
      }
      
      int64_t so_time_mono(void) {
          return (int64_t)board_uptime_ms() * 1000000;
      }
      

      What happens if the user doesn't provide an implementation? You still want the standard library to compile and work unless someone calls the missing functions. To achieve that, use weak definitions:

      // so_write_out drops the bytes and reports a full write,
      // so panic and fmt print nothing and report no error.
      __attribute__((weak)) so_int so_write_out(const uint8_t* buf, so_int size) {
          (void)buf;
          return size;
      }
      
      // so_crand_read reads no bytes. The interpretation is left to the caller.
      __attribute__((weak)) so_int so_crand_read(uint8_t* buf, so_int size) {
          (void)buf;
          (void)size;
          return 0;
      }
      
      // so_time_wall panics, because no default date is correct.
      __attribute__((weak)) so_R_i64_i32 so_time_wall(void) {
          so_panic("time: define so_time_wall for this target");
      }
      

      Now every hook gets a default, and a definition in the user code silently wins over the default one.

      Note that the defaults above behave differently on purpose. Dropping output is fine because a board with no UART (serial interface) has nowhere to print. Inventing a date is not fine, because no date would be correct.

      The same reasoning makes crypto/rand panic instead of falling back to a software generator. A "random" source that quietly returns predictable bytes would be a terrible idea:

      // crand_read fills buf with size cryptographically secure random bytes.
      // Panics if the target does not define so_crand_read.
      static inline void crand_read(uint8_t* buf, so_int size) {
          if (size <= 0) return;
          if (so_crand_read(buf, size) != size) {
              so_panic("crypto/rand: no entropy source");
          }
      }
      

      You can still use a random fallback when cryptographic security isn't needed, such as for hashing map keys or math/rand:

      // runtime_Seed returns a random 64-bit seed.
      static inline uint64_t runtime_Seed(void) {
          uint64_t seed = 0;
          // Use cryptographically secure random if available.
          if (so_crand_read((uint8_t*)&seed, 8) == 8 && seed != 0) {
              return seed;
          }
          // Fallback to deterministic xorshift64 sequence.
          // ...
      }
      

      Hosted-only

      Some things aren't worth solving with hooks, such as the os and net packages, which require a lot of target-specific code. In these cases, it's better to use a header-level guard that fails in freestanding mode:

      // so/os/os.h
      #include "so/builtin/builtin.h"
      
      #ifndef so_build_hosted
      #error "os: hosted environment required"
      #endif
      

      If user code imports os in a freestanding environment, the compiler reports an error at compile time instead of at link time or runtime.

      Testing

      "Compiles without libc" is easy to believe and easy to get wrong. The only way to be sure is to test the freestanding implementation.

      My approach in Solod is to run the freestanding packages' test suites with a WASI runtime and a small harness. The harness defines all five hooks from the Target hooks section as WASI imports:

      // ciovec is the buffer descriptor that fd_write reads.
      // The WASI ABI is 32-bit, so both fields are 32-bit.
      typedef struct {
          const uint8_t* buf;
          uint32_t len;
      } ciovec;
      
      // wasi_fd_write writes the buffers to the file descriptor
      // and stores the number of bytes written in nwritten.
      __attribute__((import_module("wasi_snapshot_preview1"), import_name("fd_write")))
      extern uint32_t wasi_fd_write(uint32_t fd, const ciovec* iovs,
                                    uint32_t iovs_len, uint32_t* nwritten);
      
      // so_write_out writes size bytes to the standard output of the WASI host.
      so_int so_write_out(const uint8_t* buf, so_int size) {
          ciovec iov = {.buf = buf, .len = (uint32_t)size};
          uint32_t written = 0;
          if (wasi_fd_write(1, &iov, 1, &written) != 0) {
              return 0;
          }
          return (so_int)written;
      }
      

      The freestanding make task builds tests from stdlib packages into a single wasm32-freestanding module and runs it with wasmtime. This covers the freestanding logic with the same tests that run in hosted mode, so no separate tests are needed.

      Final thoughts

      Here's a summary of the approach I used write a freestanding stdlib in C:

      • Choose between hosted and freestanding at compile time.
      • Use the compiler builtins when possible.
      • Implement the missing parts and port the standalone code.
      • Use explicit allocators; prefer values to pointers.
      • Declare hooks for the hardware, with weak defaults.
      • Fail fast for packages that can't work in freestanding.
      • Test in a freestanding build, not just hosted.

      I hope you find it useful too.

      If you're interested in trying this in practice, take a look at Solod's README — it has everything you need to get started. Or try it online without installing anything.

    7. 🔗 r/LocalLLaMA I just built a mini Kimi-K3 from Scratch under 250$. Already beats GPT-2 (124M)! rss

      I just built a mini Kimi-K3 from Scratch under 250$. Already beats GPT-2 (124M)! | I pre-trained a 1.02-billion-parameter on Kimi K3 replica trained on 5.00 billion decontaminated tokens for $250. This model has 1.02 billion parameters, of which 145 million are active per token. It is roughly one two-thousandth of K3 by total size. It saw 5,000,003,584 tokens, which is a rounding error against the corpora frontier models are trained on. It has never been instruction-tuned, and it has only ever done one thing: predict the next token. What it does have is K3's architecture: - Kimi Delta Attention, Gated MLA, Attention Residuals - LatentMoE with the same aux-loss-free balancer - Same activation function with the same two constants - K3's own 163,840-token tokenizer, unmodified. I report a 33.4% HellaSwag which beats the GPT-2 124M score of 28% Read the entire tutorial here: https://books.vizuara.ai/book/pretraining-a-mini-k3 submitted by /u/OtherRaisin3426
      [link] [comments]
      ---|---

    8. 🔗 smol-machines/smolvm smolvm v1.9.0 release

      What's Changed

      • macOS: name the missing hypervisor entitlement when a boot fails with -22 by @Bnjoroge1 in #948
      • Detach the API test server from the test pipe by @BinSquare in #969
      • Bring the opt-in smolfile and virtio-net suites up to date by @BinSquare in #970
      • Enforce the resolved egress policy when a machine run is served from the image cache by @BinSquare in #972
      • Repack artifact-sourced machines whose layer cache is tar-form by @BinSquare in #971
      • Delegate cgroup2 controllers at agent boot by @BinSquare in #973
      • Constrain machine names to lowercase DNS labels by @BinSquare in #975
      • Hostname containers after their machine by @BinSquare in #974
      • Write the default /etc/hosts when the image ships one with no entries by @BinSquare in #976
      • Stop egress-events from panicking when its output pipe closes early by @BinSquare in #978
      • Revert the machine-name DNS-label restriction by @BinSquare in #979
      • Fail guest connections fast when the host cannot reach the destination by @BinSquare in #981
      • Verify layer extraction reached the disk before trusting the layer cache by @BinSquare in #980
      • Let image machines without an entrypoint boot to the bare agent by @BinSquare in #985
      • Auto-shrink a machine's disk by reclaiming freed blocks to the host by @BinSquare in #984
      • Update h2 to 0.4.16 to clear the RUSTSEC-2026-0258 advisory by @BinSquare in #995
      • Fix isatty() reporting a macOS virtiofs-mounted file as a terminal by @BinSquare in #993
      • Support segmented CUDA graph capture by @BinSquare in #987
      • Add uid, gid and mode flags to machine cp and land streamed uploads inside the running workload container by @BinSquare in #997
      • Resolve guest file and socket operations in the namespace the workload actually runs in by @BinSquare in #878
      • Inject a bundled Venus Vulkan driver into GPU workload containers so stock images get Vulkan with no setup by @BinSquare in #1001
      • Use shared COW disk bases for VM-mode creates by @depombo in #991
      • Bump the workspace to 1.9.0 by @BinSquare in #1002

      Full Changelog : v1.8.3...v1.9.0

    9. 🔗 r/LocalLLaMA Qwen3.8-27b has the highest level of "agency" I've ever seen in a local model rss

      Qwen3.8-27b has the highest level of "agency" I've ever seen in a local model | Off a single prompt, given my credentials and the name of my university, qwen3.8-27b was able to successfully pull my class schedule from the kinda shitty and convoluted web of university websites. It needed no human intervention, and executed 80 tool calls. Another time, I asked it to investigate a user on a social media network, and it found one public video, downloaded it, extracted frames every few seconds so it could "watch" the video, and installed fucking openAI whisper and ran a transcription to understand the context, before selectively zooming in on and brightening some frames to see the action. That this shit is running on my own hardware (single RTX 3090) is fucking incredible, the general public doesn't realize how cyberpunk our reality already is. Quant: Unsloth's Q4_K_S (kv cache quantized to q8) Context: 150k submitted by /u/synth_mania
      [link] [comments]
      ---|---

    10. 🔗 WerWolv/ImHex Nightly Builds release

      Nightly

      53806b3 Changelog

      • impr: Add proper error dialog when a plugin fails to load due to linker issues
      • feat: Suggest not-yet-downloaded patterns in suggested patterns dialog
      • impr: Move store download API to Store API class
      • impr: Refactor Web APIs into separate files
      • git: Update more runners
      • git: Update remaining CI runners to Ubuntu 26.04
      • fix: FileAttachedData throwing when reading before writing
    11. 🔗 anthropics/claude-code v2.1.237 release

      What's changed

      • Fixed prompt caching for sessions using an LLM gateway or custom base URL
      • Added a built-in "Concise" output style: Claude leads with results and skips preamble and narration, while doing the work just as thoroughly. Select it under Output style in /config.
    12. 🔗 Rust Blog Supply chain attack on arrayref rss

      What happened

      On 2026-08-20 at 7:15 UTC we got a report that the proc-macro1 crate was malicious.

      The Rust Security Response Team verified this to be the case: the crate had a build script that was downloading a malicious payload.

      This crate proc-macro1 and others like it (proc-macro-en, aovine, arone, aronenao, tinymember) have been deleted.

      Furthermore, we discovered that the popular arrayref crate had recently been republished and made to depend on this crate, with the most recent versions yanked. We have removed the malicious version and unyanked the maliciously- yanked versions. Other crates by that author (internment, append-only- vec) were also affected so we have done the same for those, and locked the account as a precaution. We do not believe the author of arrayref to be acting maliciously, but their computer or credentials are likely compromised, and we are attempting to contact them.

      What you need to do

      We recommend you check your local dependencies to ensure these crates were not pulled in. Here are the malicious versions that we deleted from crates.io:

      • append-only-vec@0.1.9: published at 2026-08-20T07:37:49Z, deleted at 2026-08-20T09:25:24Z. Online for 107 minutes.
      • arrayref@0.3.10: published at 2026-08-20T07:15:00Z, deleted at 2026-08-20T08:41:40Z. Online for 86 minutes.
      • internment@0.8.7: published at 2026-08-20T07:34:07Z, deleted at 2026-08-20T09:04:11Z. Online for 90 minutes.
      • proc-macro1, proc-macro-en, aovine, arone, aronenao, tinymember (any versions).

      You can quickly check if these crates have been used locally by going through ~/.cargo/registry/cache with this command:

      find ~/.cargo/registry/cache -type f \( \
        -name 'append-only-vec-0.1.9.crate' -o \
        -name 'arrayref-0.3.10.crate' -o \
        -name 'internment-0.8.7.crate' -o \
        -name 'proc-macro1-*.crate' -o \
        -name 'proc-macro-en-*.crate' -o \
        -name 'aovine-*.crate' -o \
        -name 'arone-*.crate' -o \
        -name 'aronenao-*.crate' -o \
        -name 'tinymember-*.crate' \
      \) -print
      

      Thanks

      We'd like to thank the Research Team at Nextron Systems GmbH for initially discovering this and reporting it to us. We'd also like to thank Emily Albini, Manish Goregaokar, Marco Ieni, Tobias Bieniek, Ubiratan Soares, and Walter Pearce for participating in the response here.

    13. 🔗 Rust Blog Announcing Rust 1.98.0 rss

      The Rust team is happy to announce a new version of Rust, 1.98.0. Rust is a programming language empowering everyone to build reliable and efficient software.

      If you have a previous version of Rust installed via rustup, you can get 1.98.0 with:

      $ rustup update stable
      

      If you don't have it already, you can get rustup from the appropriate page on our website, and check out the detailed release notes for 1.98.0.

      If you'd like to help us out by testing future releases, you might consider updating locally to use the beta channel (rustup default beta) or the nightly channel (rustup default nightly). Please report any bugs you might come across!

      What's in 1.98.0 stable

      Algebraic floating-point methods

      The floating-point types f32 and f64 now have "algebraic" methods for addition, subtraction, multiplication, division, and remainder. These allow optimizations on these operations using the algebraic properties of real numbers, even though these properties do not hold with the limitations of floating-point representations. The exact set of optimizations is not specified, but may be similar to the kind of optimization you would see with the -ffast-math option in other languages.

      For example, floating-point addition is not associative, so a sum like a + b + c + d must be evaluated in the left-associative order in which it is parsed, like ((a + b) + c) + d. If you write the same sum as a chain of algebraic_add calls, then the compiler is free to reorder it, perhaps like (a + b) + (c + d) to evaluate the partial sums simultaneously. Broader loop-vectorization is often enabled by using these algebraic methods as well.

      These methods are non-deterministic, since the compiler is free to choose different optimizations, but they never cause undefined behavior. See the library documentation and the original API change proposal for more details.

      Buffered integer formatting

      All of the primitive integer types now have a format_into method that takes a &mut NumBuffer<Self> parameter, which is a buffer that is large enough to hold the decimal format of any value of that type. The buffer itself is opaque, but the method returns the formatted &str with a lifetime borrowed from that buffer.

      This method also bypasses much of the dynamic dispatch that you would get with buffered write! formatting, which can be a boon to performance. The itoa- benchmark repo now shows that format_into performs similarly to itoa itself, so this could serve as a standard replacement for that dependency and others like it.

      Fix interaction between ManuallyDrop and Box

      Prior to Rust 1.96.0, there was a bug in the Rust compiler, which made the following code undefined behavior:

      let mut x = ManuallyDrop::new(Box::new(1));
      unsafe { ManuallyDrop::drop(&mut x) }
      let x = x; // UB!
      

      This is because the compiler considers it undefined behavior to move a Box that has been dropped (deallocated), and ManuallyDrop used to propagate that, such that moving ManuallyDrop<Box<_>> where the box has been dropped would also be considered UB.

      In Rust 1.96.0 we fixed this, so this code was no longer UB. In this release we have updated the ManuallyDrop documentation, providing a stable guarantee that this code will continue to not be UB in the future. See ManuallyDrop docs and the related RFC 3336 for more information.

      Stabilized APIs

      Other changes

      Check out everything that changed in Rust, Cargo, and Clippy.

      Contributors to 1.98.0

      Many people came together to create Rust 1.98.0. We couldn't have done it without all of you. Thanks!

    14. 🔗 Console.dev newsletter TanStack Table v9 rss

      Description: Headless tables.

      What we like: Provides underlying functionality for sorting, paging, selection, spans, pinning, columns, and filtering data grid UIs. You maintain control of styles and interactions. Integrates into various frameworks e.g. React, Vue, Solid, Svelte. Lots of focus on performance.

      What we dislike: Potential bundle size increased 14 kB in v8 to 25 kB in v9, but tree shaking will help reduce that to just the required features.

    15. 🔗 Console.dev newsletter TurboVec rss

      Description: Rust vector index.

      What we like: Built on Google’s TurboQuant algorithm. Add vectors without any training step or tuning. Bundles optimized kernels for various architectures. Runs locally and self-hosted. Integrates into common frameworks (LangChain, LlamaIndex, Haystack, Agno).

      What we dislike: Good integration docs, but minimal operational docs about how to run it.

  2. August 19, 2026
    1. 🔗 IDA Plugin Updates IDA Plugin Updates on 2026-08-19 rss

      IDA Plugin Updates on 2026-08-19

      New Releases:

      Activity:

    2. 🔗 PrimeIntellect-ai/prime-agent v0.7.4 release
      • Fixed model searches ranking stronger matches ahead of weaker signed-in matches while preferring signed-in providers for equivalent results (#539 by @eliebak).
      • Fixed large IPython variables repeatedly slowing later turns by excluding them from persistent snapshots and removing them when context is compacted.
      • Fixed daemon socket paths being used verbatim in identity derivations: on supported platforms, --daemon-socket spellings differing only by duplicate or trailing slashes now normalize to one canonical path, so worker-descriptor namespaces, daemon log files, and persisted descriptors agree.
      • Added a thinking option to rlm.run for spawning subagents with an explicit reasoning level; invalid levels for the resolved child model fail spawn.
      • Changed opening the agents view (full or scoped) with a draft prompt to auto-stash the draft instead of refusing; the draft is restored into the editor when the session is reopened.
      • Fixed Shift+Enter no longer inserting a newline in terminals that send a literal \n (for example a Ghostty shift+enter=text:\n mapping): the byte decoded as ctrl+j and triggered the new edit-diff toggle instead of the editor newline.
      • Removed a system prompt paragraph referring to an async bash() kernel helper and managed jobs that do not exist in the runtime.
      • Changed RLM guidance to orchestrate independent workers in parallel, use available async shell helpers safely, end the turn instead of sleeping, polling, or blocking on long awaits, provide proactive outcome-focused progress updates from root agents, and use simplified technical English for user-facing prose.
      • Fixed new top-level daemon sessions inheriting an RLM child depth from the supervisor process.
      • Fixed active goals stalling after a mid-goal automatic compaction when the previous continuation prompt was already running: only undelivered continuations deduplicate, so a fresh continuation is queued instead of being suppressed.
    3. 🔗 Simon Willison Conceptual integrity and counting lines of code rss

      Last week I recorded an episode of the Talking Postgres podcast with Claire Giordano on the subject of "How AI is changing software development". We had a really great conversation. Here are a couple of my highlights from a lightly edited transcript (prompt to Claude: "very minor edits to remove disfluencies").

      This is the latest version of an argument I've been trying to build about why sometimes it does make sense to talk about lines of code as an indicator of productivity with coding agents, at 35:01:

      A lot of people will tell you it makes no sense to measure productivity in lines of code. I’d actually disagree, because there’s a hard limit. In the before-times, a software engineer could produce a few hundred lines of production-ready code per day — and 200 lines of working, debugged, production-level code is an incredibly good day. Most days you’d produce 50 or 60.

      If agents let you produce a thousand lines of debugged code, that really is a very meaningful improvement — as long as the code is the same quality: maintainable, tested, all of that. You can get to that point with agents, but it takes a huge amount of skill and knowledge and experience. That’s what senior engineers are made of.

      I can do way more work as a single engineer than I could without agents. So you could argue, why should a company have more than one engineer? Beyond the obvious bus factor thing — a team of one is a very badly designed team — the answer is that the new limiting factor is cognitive capacity. I can churn out code a hundred times faster. I don’t have the cognitive capacity to stay on top of 100 times the amount of code. So you still need a team of engineers, so you can load balance that cognitive capacity across the team.

      And this section on conceptual integrity at 46:03, which Claire equated to the Winchester Mystery House!

      Simon: There’s a concept in The Mythical Man-Month — conceptual integrity — where well-designed software has an integrity to it: there are no surprises in it, it covers exactly the right domain of things, everything fits together and makes sense. That’s so much harder with coding agents, where you can have an idea for a feature, run a prompt, and five minuteslater you’ve got the feature. Your software grows little weird bumps in funny different directions.

      Claire: You know my analogy for that? The Winchester Mystery House.

      Simon: It’s got 140 rooms, because the woman who built it was the widow of the guy who invented the Winchester rifle, and her psychic told her she’d be haunted by the ghosts of everyone killed with that rifle unless she kept building the house forever. So for 40 years she kept adding new rooms. That’s exactly the problem with coding agents and software: it’s very easy to keep adding new rooms, because the cost of adding those rooms is so much cheaper. What you end up with is something where the conceptual integrity falls apart — and then it’s harder to make decisions about it.

      It all keeps coming back to discipline. It used to be that the discipline was enforced on you by the amount of time it took. You’d come up with an idea for a crazy feature and think “yeah, but that would take me a week — I cannot justify that, so I’ll forget about it.” If it takes an hour, it’s so much easier to justify.

      (Side-note: the Wikipedia article includes credible sources that dispute the story about the psychic.)

      You are only seeing the long-form articles from my blog. Subscribe to /atom/everything/ to get all of my posts, or take a look at my other subscription options.

    4. 🔗 HexRaysSA/plugin-repository commits sync plugin-repository.json rss
      sync plugin-repository.json
      
      No plugin changes detected
      
    5. 🔗 anthropics/claude-code v2.1.236 release

      What's changed

      • Added ANTHROPIC_DEFAULT_MODEL environment variable: sets the model new sessions start on, while a /model pick still overrides it and persists across restarts (unlike ANTHROPIC_MODEL)
      • Added notify_when_idle to cross-session SendMessage: ask another Claude Code session on this machine to send one notice when it next goes idle — opt-in, one-shot, no polling (macOS and Linux)
      • Sandbox: on macOS, wildcard read-deny rules (e.g. **/.env) now take precedence inside allowed read regions, cover matched directories' contents, and can't be bypassed by renaming the denied file
      • Fixed clipboard copy, background housekeeping, background sessions, and local MCP logs breaking after the directory a session had switched into was removed (since 2.1.229)
      • Fixed the fullscreen renderer failing permanently after a single failed start: it now falls back to the classic renderer instead of exiting on every subsequent launch
      • Fixed the /model picker rendering taller than the terminal: it now shows only as many models as fit the window, with the rest reachable by scrolling
      • Fixed SendMessage calls being rejected when a malformed closing tag left the message text inside the summary field
      • Fixed unhandled promise rejections when a subprocess fails to start, for example powershell.exe on WSL with Windows interop disabled (regression in 2.1.234)
      • Fixed fullscreen mode sometimes not showing a newly sent message until the next update after the terminal was resized
      • Fixed a blank band that could remain above the prompt after clearing a multi-line prompt, and panes not repainting after resizing the terminal away and back, in fullscreen mode
      • Fixed the managed-settings approval prompt sometimes not appearing at startup while still capturing the first keypress as approval
      • Fixed terminal tab titles jumping in tmux (iTerm tmux integration): the title is now written only when its text changes instead of animating every 960ms
      • Fixed an unclear error when the cloud environments list came back empty or malformed
      • Fixed the Fable 5 first-time usage-credits prompt auto-selecting the fallback model after 60 seconds with no answer when using Remote Control
      • Fixed spinner tips never appearing, with a repeated background error, when the cached guest-pass reward in ~/.claude.json was malformed
      • Fixed skills hot-reload in SDK/VS Code sessions raising an error on every skills change after the session's working directory was deleted (2.1.229+)
      • Fixed self-hosted runner sessions released on idle, retire, or startup timeout occasionally resuming on another runner before the post-session hook had finished
      • Fixed the Clawd mascot's eyes and feet rendering unevenly in iTerm2 at some font sizes
      • Fixed occasional runaway session recaps: recap text (automatic and /recap) is now capped at 400 characters, cut at a word boundary
      • Improved startup performance: the session counter is now written in the background
      • Improved auto mode: Monitor allow rules are now set aside while auto mode is active, so Monitor commands are reviewed the same way Bash commands are
      • Improved auto mode on Bedrock, Vertex AI, and Foundry, and when telemetry is disabled: the classifier now uses the same defaults as on the Claude API, including severity-scored classification
      • Improved auto mode: the git status check can no longer be fooled by a repo's status.showUntrackedFiles=no setting into reporting a clean tree
      • Changed the /model picker to highlight only the newest model's name, so the highlight marks the new release rather than an arbitrary subset of the list
      • /goal: an idle session whose goal is parked behind long-running background work now checks in automatically after 30 minutes (then 1h, 2h) instead of waiting for you to return
      • /usage now shows the usage-credits spend row for Team and Enterprise members, and shows a capped row at 0% before anything is spent
      • SIGTERM in print/SDK mode no longer records an interrupted turn or synthetic tool denials before exiting; running commands are still terminated and the process still exits with code 143
      • Pressing Enter on a slash-command typo or a command unavailable in this session now reports it instead of running the closest fuzzy match; prefixes and aliases still run
      • Remote Control now marks a session offline within seconds when the CLI exits or its terminal closes
      • SendMessage now refuses further messages to a session up front once a rapid burst would exceed what that session's inbox accepts, instead of reporting them sent while they were dropped
      • Aligned the session title chip on the prompt border with the footer's right edge
      • Right-aligned footer items (goal indicator, session state, background agent status) and truncated notices now share a consistent right margin with the rest of the prompt area
      • [VSCode] Added screen reader support for the transcript: live announcements for replies, permission requests, errors, and status changes, plus per-turn heading navigation
    6. 🔗 r/LocalLLaMA Introducing Qwen3.8-27B Dynamic v3 Unsloth GGUFs rss

      Introducing Qwen3.8-27B Dynamic v3 Unsloth GGUFs | Hey everyone! We’re releasing new Qwen3.8-27B GGUFs with 10% higher accuracy for the same size. This uses a new version of Dynamic v3.0 Unsloth Dynamic V3 outperforms others by >10% on Div-300, KLD & more benchmarks. We also release 1-bit quants that retain 77% accuracy. Run on 8GB RAM. Some of you already saw we updated our quants a few hours ago. No, nothing was broken, nothing needed fixes (I don't know why people even said this since it's a complete fabricated story). This was purely an update to make them EVEN BETTER. We do not train on the imatrix calibration dataset, and we do NOT use QAT or QAD. Everything is done through post-training quantization. Our imatrix file used is available for the community to test, evaluate, and use. We encourage researchers and developers to create variations and fine-tunes of Qwen3.8 using our Unsloth quants/imatrix. You can read our over fitting analysis as well. Blog with all details and more benchmarks: https://unsloth.ai/docs/basics/dynamic-3.0-ggufs GGUF: https://huggingface.co/unsloth/Qwen3.8-27B-GGUF Enjoy! We also will be doing a new Unsloth Desktop update today: https://github.com/unslothai/unsloth We had A LOT of updates and will be introducing auto compaction, allowing external APIs to do tool calling and more. submitted by /u/danielhanchen
      [link] [comments]
      ---|---

    7. 🔗 The Pragmatic Engineer The Pulse: Grok’s CLI caught uploading all your local files to the cloud rss

      The Pulse: Grok’s CLI caught uploading all your local files to the
cloud

      Hi, this is Gergely with a bonus, free issue of the Pragmatic Engineer Newsletter. In every issue, I cover Big Tech and startups through the lens of senior engineers and engineering leaders. Today, we cover one out of four topics from a previous The Pulse issue. Full subscribers received the article below four weeks ago. If you 've been forwarded this email, you can subscribe here .

      Last week, xAI (Elon Musk's AI company, now part of SpaceX) released the Grok 4.5 model, built by Cursor (an acquisition), trained on SpaceX GPUs, and branded "Grok." It's a pretty good model; benchmarking as close in coding capability to Opus 4.8 and GPT 5.5 - while being 60-70% lower cost.

      Grok 4.5 can be used via API, but is easiest used via the Grok Build coding CLI. So, that's what many devs did. Some of them have noticed something really weird: the CLI is uploading all their local files in their working directories! Here's an independent AI safety researcher known as 'Cerblab' documenting what is happening (emphasis mine):

      "xAI's official Grok Build coding CLI (grok), on a normal consumer login, does three things worth documenting precisely:It transmits the contents of files it reads -- including a .env secrets file -- to xAI, verbatim and unredacted. The secret appears in two channels: the live model turn (POST /v1/responses) and a session_state archive uploaded and accepted (HTTP 200) via POST /v1/storage -- the endpoint the binary routes to the grok-code- session-traces GCS bucket (see section 5).It uploads the whole repository -- every tracked file's content plus git history -- independent of what the agent reads. Grok packages the workspace and uploads it via POST /v1/storage. Proven directly: on a real codebase, with the prompt "reply OK, do not read any files", Grok uploaded the entire repo as a git bundle (POST /v1/storage -> 200); git cloning the captured bundle recovers a file the agent was told not to open -- src/_probe/never_read_canary.txt -- with its unique marker verbatim, plus the full git history (appendix uploaded_repo.bundle). And it scales: on a 12 GB repo of never-read random files, /v1/storage moved 5.10 GiB, all HTTP 200 (truncated mid-stream), while the model-turn channel moved just 192 KB -- a ~27,800× ratio that pins the upload to the codebase, not to what was read. No storage upload failed; the only non-200s were a model-usage quota (402/429) on /v1/responses and one unrelated 404 -- not a storage size cap.The storage destination is a Google Cloud Storage bucket, grok-code-session-traces (not AWS S3) -- named verbatim in the binary and in a captured metadata.json (gs://grok- code-session-traces/…). I did not find this mechanism surfaced in the CLI's install/quickstart materials (not an exhaustive docs audit -- §7), it is active by default, and disabling "Improve the model" does not turn it off (/v1/settings still returned trace_upload_enabled: true; §6).

      None of this proves xAI trains on the data -- that is a policy question addressed in [section] 6. What is proven is transmission, acceptance, and storage."

      There's so much wrong with this approach! To name a few:

      • Sending your codebase over the context window is not normal. All AI agents send over their context window to the server that runs the LLM. That's where tokenization happens and the context is appended to the session. This means that other AI agents send over some part of the code that they read in their context window.
      • No need to send over the source code to index it. Indexing the codebase is important for efficient code lookup, and we covered how Cursor does this in a privacy-conscious way by indexing a user's codebase locally, creating embeddings, and sending those embeddings to the server. Cursor's server does not store any of the user's codebase though. Except that Grok CLI transferred all users' codebases to a cloud bucket! Given Cursor and Grok are now combined as part of SpaceX, it 's a real head scratcher why the Grok CLI isn't doing what Cursor always has.
      • Sending over unencrypted .env files is reckless. Local .env files store secrets, database access tokens, service access tokens, and more. These are sensitive pieces of information that need to be handled with care. If transmitted, they should be encrypted at the very least. Grok / SpaceX storing them on the GCP storage bucket - likely unencrypted - is flat-out unacceptable and reason enough for any sensible company to ban usage of Grok CLI.
      • No good reason to upload git history. Sure, seeing Git history could be helpful when training an AI model.
      • It 's malicious to not tell devs anything about it. Developers using Grok CLI haven't been asked to opt into this data upload, nor notified of it. Most evidently had no idea this has been happening, and are understandably furious after Cerblab's writeup went viral.

      SpaceX throws devs "under the bus"

      Caught red-handed, Grok CLI disabled file uploads with a remote feature flag. AWS engineer Wes Eklund started tracking upload functionality with the CLI, finding that the uploads suddenly stopped due to a feature flag being flipped by the Grok team, pausing data collection.

      But the code functionality to stream all local files to the server, unencrypted, remained present in the CLI, even in later updates. Read an in-depth analysis by Wes.

      SpaceX 's official response was pretty laughable, not explaining why .env files and .git history were uploaded, and adding that enterprise customers with zero data retention (ZDR) enabled were the only ones unaffected by underhanded, secret uploading of users' local files:

      The Pulse: Grok’s CLI caught uploading all your local files to the
cloudTranslation: "Enterprise users with ZDR turned on were not impacted. To everyone else: we didn't tell you about this hidden/privacy command, but it's your fault. Source: SpaceX

      My initial reaction is what a condescending response by SpaceX, swiftly followed by the question: does Grok CLI even have enterprise customers? It might be few, given how reckless the team evidently is by uploading unencrypted secrets! SpaceX CEO Elon Musk chimed in with a post that seemed to be almost trolling angry users. He wrote:

      "SpaceX policy regarding data retention.

      It is actually helpful for debugging issues if we can retain some amount of data, so allowing this would be appreciated, but your privacy settings are always respected."

      This makes it worse because SpaceX has been secretly uploading far more data than is "useful for debugging!" Uploading the git history and sensitive .env files is not , in any way, useful for debugging. Also, Grok/SpaceX did not upload "some amount of data"; it uploaded every last file it could find in your local folder.

      The developer community is justifiably upset to read SpaceX and Musk pretending that Grok has only uploaded scraps of data purely for debugging purposes. This time, even fans of SpaceX and Grok are speaking out against Musk and his company's behavior. AWS engineer Wes Eklund:

      "Elon, firstly, huge fan of everything you work on. Completely understand the need for some trace data to improve customer experiences.

      From what I've researched, it seems to be much more than just trace debugging issues. It seems to be entire code repos with sensitive information just collected entirely.

      Your Google Cloud blob storage must have petabytes of code repos from us.

      Not ideal."

      Sam Altman pushes Grok to open source Grok CLI

      OpenAI CEO, Sam Altman, also posted, using a term Musk often employs when commenting on things he disapproves of in society, and hinting at the benefits of Codex which doesn't upload your whole local filesystem, or mess with .env files and the git history. The Codex harness is open source, so secretive file-upload functionality would be visible in the source code:

      The Pulse: Grok’s CLI caught uploading all your local files to the
cloudAltman uses Musk 's trademark "concerning" remark against him. Source: Sam Altman

      Musk clearly read it and responded a few hours later:

      The Pulse: Grok’s CLI caught uploading all your local files to the
cloudMusk committing to open sourcing Grok CLI. Source: Elon Musk

      A day later, (15 July), SpaceX did indeed open source Grok CLI. Altman's comment seemingly hit home. Also, SpaceX has stated it is deleting data from its servers, writing:

      "We disabled default retention for all Grok Build users starting on July 12th. Additionally, we are deleting all coding data that was previously retained, ensuring every user's preferences are respected. With these steps, Grok Build goes beyond other major coding products to protect user privacy."

      The open sourced repo is rushed, unsurprisingly. Kernel engineer Elliot Arledge used the repo and reported what he found:

      "Out of the box, cargo test --workspace doesn't compile. 190+ errors, all one bug class: cross-crate test helpers hidden behind #[cfg(test)], which Bazel's per-target test builds tolerate but Cargo doesn't, because a dependency is never compiled in test mode. The default-bazel feature is declared in ~20 manifests and wired to nothing. Ungating the benign helpers and putting the signing-key test seam behind a dev-dependency-only feature gets the suite running for the first time on the published tree: 24,663 passed, 28 failed, and every one of the 28 is a pre-existing bug the broken build had been hiding (tests reading your real ~/.claude/settings.json, a tool missing from the registry, macOS /var symlink breakage, one Theme::current() race).

      Questions for the team:

      Did anyone try downloading this repo and running it before publishing?"

      In fairness, it's clear from the outside that the dev team was instructed to open source the repo ASAP, and did just that. I would expect improvements to allow the developer community to compile the repo and run tests will come later.

      On a side note, I wonder what the mood within Grok is. A mandate to open source the product - that was not built to be open sourced! - and doing it in a couple of hours hints what it's like to work there. Some folks doubtless would find it thrilling and a big challenge, while others likely find it pretty stressful having the "Eye of Sauron" (the CEO) upon them!

      Can companies trust Grok now?

      To hand it to Grok/SpaceX, the last 24 hours of this incident saw some impressive execution. In contrast, everything that took place beforehand screams "amateur hour":

      • Did no one in teams that wrote the functionality to upload the complete local codebase raise concerns that this would be unacceptable to devs?
      • How did uploading secrets without encryption not raise alarm bells?
      • Why greenlight any of it without opt-in and do it in a secretive manner?

      For sensible companies generating revenue with software, vendors who are allowed to access their codebase are limited to those that can be trusted. Grok has demonstrated it is unprepared to handle codebases with security fundamentals in mind.

      There's now frantic backpedaling and rapid open sourcing, but my impression is that this is simply because Grok/SpaceX was caught red-handed, and understands the risk of losing enterprise contracts caused by this flagrant breach of trust. This is why SpaceX's communication mentioned that enterprise customers with zero data retention have not been impacted!

      Trust is earned in drops and lost in buckets; SpaceX/Grok will likely learn this. By secretly uploading codebases and secrets, Grok CLI revealed itself as an untrustworthy coding agent that's a risk to use. Meanwhile, all the major competitors - Claude Code, Codex, OpenCode, Gemini CLI - have never violated user trust this way.

      Grok CLI can rebuild trust, but it'll likely take years of no security-related incidents to prove they are an open source-first product (that they were not, just a day ago!), and also demonstrate that they care about "normal" developers, and not just enterprise clients with ZDR turned on.

      This incident is a good reminder of why planning and process can slow shipping speed, but increase revenue generation. I would wager that the Grok team has shrunk SpaceX's enterprise subscription prospects for the foreseeable, in the name of saving a few hours on security reviews, learning what other coding harnesses do with codebase uploads, or even just asking the Cursor team!

      I predict Grok/SpaceX will have to offer very high usage limits inside the Grok CLI to convince devs to take a risk on running this software on their system. And they will have to undercut OpenAI and Anthropic API pricing massively for any security team to greenlight use of a CLI that just last week was sending .env secrets unencrypted to their GCP buckets.

      Of course, SpaceX/Grok will be just fine as it has the capital to fix things. It will now just be a lot more expensive and time-consuming to fix something that was likely caused by a few engineers wanting to make debugging easier!

      Read the _full _The Pulse issue__ , or check out this week 's The Pulse . The full issue additionally covers:

      1. New trend: concern about massive increase in code review load. Top of mind for engineering leaders: what to do about the ever-growing code review load, and how devs are starting to review code less thoroughly than before? Many questions, but few proven solutions. Send comments about what you see working.
      2. Are more devs at enterprises upset about enterprise pricing by AI labs - and does it matter? I got a message from a reader baffled to learn their company pays 20-30x the price for tokens than their own $20/month Claude Code / Codex subscription. It may show how valuable AI coding tools are.
      3. Linux creator: AI "clearly useful." Inside the Linux kernel maintainers group, the discussion veered onto whether Linux should consider banning AI contributions, similar to how some FOSS projects have done so. Linus Torvalds weighed in and made it clear that AI is useful, everyone should decide whether to use it, but no one is allowed to tell others what tools to use. Given AI is an increasingly capable tool, it would be foolish to not use it as such.

      Read the full The Pulse

    8. 🔗 @binaryninja@infosec.exchange Sidekick can use the terminal now! With the new run_terminal_command tool, mastodon

      Sidekick can use the terminal now! With the new run_terminal_command tool, agents can jump into bash or PowerShell right alongside their Binary Ninja analysis. Unpack an archive, run binwalk, poke through git history, curl a URL, or inspect files that never even made it into Binary Ninja. Commands still require your approval by default. https://sidekick.binary.ninja/blog/sidekick-26-1-a-proper-home-for- sidekick/#sidekick-can-use-the- terminal

    9. 🔗 MetaBrainz Growing pains: An update on the ListenBrainz service status rss

      TL;DR: The growth of ListenBrainz has caught up with us, and our limited team is working on replacing central parts of our infrastructure. In the meantime, many features are unstable.

      We are victims of our success. While in the long term this is a good problem to have, the sharp increase in users over the past year and a half has left our infrastructure cracking at the seams.

      The good news: Your listens are being stored and imported without issue, even if there are delays. The core service is not compromised and we ask that you please keep submitting your listens. Your stats and playlists will be back!

      However, we know that every other week your statistics, weekly playlists, and other features fail to generate for everyone, and cause crashes. We know it is frustrating, and we share that feeling.

      While we are aware of the issues, fixing them is far from simple and requires us to completely rework all the crucial parts of our infrastructure.

      In the interest of transparency, here are our main issues and what we are doing to fix them:

      How many listens?

      Our database dumps have become too big for us to process.
      We went from 0 to 1 billion listens in 7 years - then to 2.5 billion in the next one and a half years.

      Generating and copying our (now huge) database dumps causes crashes as our limited servers run out of memory and disk space.

      We are working on improvements, but for each change we need to wait two days to be able to run tests.

      In addition, we are moving away from TimescaleDB, a Postgres database extension for time-series data, in favour of vanilla Postgres with a new partition scheme.

      We found that TimescaleDB was not adapted for our use case of working with historical imports or deleting listens and users (and all their listens), as well as large gaps between listens, all causing some very slow queries.

      Where are my stats, goddammit?

      Our statistics and playlist calculation infrastructure, a Spark cluster of 5 servers, is running out of memory and crashing, from one task or another. This used to happen once every few months but is now a weekly occurrence.

      This is the issue which is breaking stats, playlists, user similarity, unlinked listens, fresh releases and more.

      We are moving to using Clickhouse instead for all statistics calculations, which will free up the Spark cluster to be used for generating playlists and other tasks.
      This move is taking some time, as a single person in the team carries the responsibility of rewriting essential code and testing everything carefully.

      This will eventually open the door to requesting stats for an arbitrary time range instead of being limited to this and last week/month/year, a hotly requested feature that is not possible with our current system.

      The scraping situation

      To make matters worse, the entire internet is being bombarded by unscrupulous bad actors (looking at you, AI companies) that don't follow the rules and try very, very hard to evade any measure meant to limit them.

      They rent botnets of millions of residential IPs so they can scrape our APIs and websites incessantly, over and over again, while evading detection, all for data that they could download for free.

      They cause surges of 5x the usual traffic across all our projects and slow everything down on our resource-constrained infrastructure.
      It also forces us to spend time dealing with these DDOS- like surges instead of working on our other pressing issues.

      For ListenBrainz specifically, we have had to disable some features/endpoints that during scraping waves made the website completely unreachable for everybody.

      But wait, there's more…

      The loss of our founder in late February was big blow to our team.

      Rob was one of the custodians of Listenbrainz infrastructure, but also a central ListenBrainz team member.

      We have had to cross-train our ListenBrainz dev team -it is only three of us- to deal with infrastructure and other new aspects, as well as reorganize priorities to deal with the day-to-day operations while the foundation was in the process of hiring a new executive director.

      We have been so greatful for the wonderful patience and kindness shown to us by you, our users and community, as we work through these growing pains. Keep submitting listens and let's grow together!

      • your ListenBrainz Team
    10. 🔗 Armin Ronacher What Is Reasoning rss

      A few weeks ago a paper was shared that showed how to extract reasoning traces from closed-weight models. Together with online discussions about tricking models into leaking them, it made me investigate it more out of curiosity. Twitter seems full of half-truths and confusion about how this works, so perhaps this helps some to understand what is happening.

      Hiding Traces

      Reasoning traces are usually hidden from us. We have lamented this, but mostly have to accept it. Open-weight models thankfully reveal them, and from their behavior you can see that their traces can be long and confusing. This is probably a good reason to separate them from what is normally shown to users.

      At minimum, UIs need to detect them. The industry has done a good job at making reasoning traces sound special and exotic, but they really are just text: the model is trained to emit its thinking into a scratchpad as part of its response, before its final answer.

      GPT-OSS's Harmony response format makes this easy to see:

      <|channel|>analysis<|message|>
      I need to work this out ...
      <|end|><|start|>assistant<|channel|>final<|message|>
      The answer is ...
      <|return|>
      

      The markers are special tokens, but the reasoning between them uses "the same text" as the final answer (just that GPT chain-of-thought text sounds really funny). When the model samples the analysis channel token, a parser routes the following text into a separate stream exposed through the Responses API. For closed models, presumably a simple model redacts and summarizes it.

      Reasoning Effort

      How much budget goes to reasoning? Earlier APIs exposed reasoning token budgets, making it seem like a property of the sampling process. In reality, reasoning effort is baked into the system prompt. GPT-OSS puts this into the system prompt:

      Reasoning: low
      

      That's it. Training produces the resulting behavior, such as emitting the token sequence that switches to the analysis channel. This also explains why changing the effort invalidates the KV cache. I think closed GPT models call reasoning effort "juice," since you can ask most models how much juice they have.

      In DwarfStar for DeepSeek with max reasoning this is added to the system prompt:

      Reasoning Effort: Absolute maximum with no shortcuts permitted.
      You MUST be very thorough in your thinking and comprehensively decompose the
      problem to resolve the root cause, rigorously stress-testing your logic against
      all potential paths, edge cases, and adversarial scenarios.
      

      Don't Think

      The destination of reasoning tokens is therefore a learned convention: the model is trained to keep scratch work out of the final channel. Trick it into thinking it is in that channel and it may leak tokens. We have even seen older models, when thinking is disabled, reason into the bash tool and echo their thoughts to /dev/null.

      So in some sense the only "special" behavior for some models is not to think. That at times is done by "mechanically" removing the model's usual ways to think. In DwarfStar, disabled thinking uses the prefill </think>, while enabled thinking uses <think>, which are the tokens that close and start thinking. GPT-OSS doesn't prefill but lets the model decide either way on its own.

      But presumably, some inference APIs prefill the opening token when reasoning is enabled, so the model never samples it itself and might prevent the sampling of the reasoning token when disabled since it can be trivially detected. This may explain why a custom think tool can trick models into putting some reasoning where it should not go — but only when native reasoning is disabled.

      Fun fact: this blog post triggered safey checks

      Hilariously enough I was unable to use GPT 5.6 terra for spell and grammar checking on this blog post because of safety filters. Had to switch to Kimi.

      GPT-5.6-terra refusing to spell-check this blog
post

    11. 🔗 Ampcode News MCP in Orbs rss

      You can now connect remote MCP servers on ampcode.com and use them in orbs, the TUI, and with Puck.

      The MCP settings screen showing preconfigured MCP servers

      To add an MCP server:

      1. Visit ampcode.com/settings/mcp-servers,
      2. Select a pre-configured server or click "Add MCP Server"
      3. Log into the server using OAuth,
      4. Start a thread in an orb, the TUI, a runner, or talk to Puck:

      Amp supports connecting hosted MCP servers using Streamable HTTP with OAuth or Bearer tokens, personal and workspace configuration of MCP servers, and the use of MCP-provided tools.

      MCP Apps, Resources, and Prompts are not supported.

    12. 🔗 Ampcode News Pass the Orb to the Left Hand Side rss

      You can now bring your team into your orb.

      Use @ to tag members of your team, and they'll be able to view, drive, and send chat messages to the thread.

      We've been having a lot of fun with this feature recently. Here are a few examples:

      Thorsten Ball passes a request for an orb sticker from Tim
      Tim Lucas mentions Lewis Metcalf to ask for feedback on a code diff
      A teammate mentions Tim to ask for a documentation review
      A message copies Camden Cheek into a thread
      A teammate mentions Rocko Reager to ask for help with a sidebar bug
      Tim Lucas mentions Rocko Reager to ask if he saw a message
      Teammates coordinate synchronized orb animations in a thread
      Tim Culverhouse mentions Lewis Metcalf before asking Amp to start work
      Thorsten Ball mentions Brett to ask about shipping an orb navigation change
  3. August 18, 2026
    1. 🔗 IDA Plugin Updates IDA Plugin Updates on 2026-08-18 rss

      IDA Plugin Updates on 2026-08-18

      New Releases:

      Activity:

      • claude-marketplace
      • diaphora
        • f997fd34: Update README.md
        • 31282dee: Merge branch 'master' of ssh://github.com/joxeankoret/diaphora
      • disrobe
        • ebdb7657: sleigh: build the unsupported effect name from a sanitised mnemonic s…
        • 94e0f342: binfmt: recover stuffit 5 entry dates, type, creator and finder flags
        • 7df8d145: dotnet: record how a nativeaot fixture project must be named before i…
        • 0cd5a4d2: cli: grade rar recovery through auto at one and four workers
        • d63e8083: binfmt: recognise stuffit 5 by its own signature and extract its memb…
        • f45554bf: binfmt: emit rar members as container chain children
        • bc019223: dotnet: record whether a nativeaot signature came from managed metada…
        • d4fcd174: binfmt: track the rar fixtures in the ignore list and correct the stu…
        • 557408e2: hermes: recover labelled loop exits and read the one-byte string key tag
        • afd8c25e: binfmt: recover rar 2.9/3.x filter records and mixed lz and ppmd blocks
        • eec74559: binfmt: decode stuffit method 5 forks and stuffit 5 containers on the…
        • da660ddd: arm32: choose the decode mode from the elf mapping symbols instead of…
        • 1f318d00: native: keep aarch64 d-register classification off callee-saved spill…
        • ef62de45: native: exercise the aarch64 simd memory forms on the decompile comma…
        • 20e738af: binfmt: decode lzh pm1 and pm2 members and count the aarch64 nan grad…
        • 4ff291f9: sleigh: vendor the aarch64 simd specification whole and lift the rema…
        • 6c9d4db9: native: recover aarch64 stack floating parameters and grade thirty sc…
        • 30e1bdbe: py: keep the guard on a with region a conditional encloses and republ…
        • dcf62039: recovery: lift aarch64 scalar floating transfers and reattach void an…
        • 445c68ce: native: restore four aarch64 recoveries and report a non-equivalent p…
      • distro
        • a20f55f4: restrict-egress 0.10.1: version bump for the diagnose allow-set fix
        • 29efc0cf: restrict-egress: read allow sets after resolving in diagnose
        • e90be796: restrict-egress 0.10.0: couple DNS answers to the allow set; add diag…
      • eject_idb
        • 0027b134: ci: bundle the standalone signaler binaries into a single tools zip
        • ee90498e: Maintenance release (v0.0.4)
      • ida-free-mcp
        • fb3b0c89: fix(decompile): return the requested function, not a stale pseudocode…
      • IDA-PRO-MCP
        • 0211e84b: chore: add repository URLs and author metadata
        • f6b1001c: Initial commit: IDA Pro MCP server with 87 tools
      • IDAPluginList
        • cc791482: chore: Auto update IDA plugins (Updated: 19, Cloned: 0, Failed: 0)
      • mytools
        • bb7ec0c6: Update rootfs-tools skill and macOS workflow
      • twdll
        • 4e80db7c: feat(attila): add ARTSET_SCRIPT_INTERFACE and character portrait API
        • ce1aac80: docs(cai): create Campaign AI Telemetry and Decision Intelligence Arc…
        • 3a7ebe40: fix(cai): use exact DWARF enum names for OCCUPATION_DECISION
        • 2ebaede2: feat(cai): add Campaign AI occupation telemetry and decision logging
        • 5915a726: feat: add region religion breakdown api and standardize list method n…
        • 5a7bade4: fix: preserve proportional health and sync bodyguard in ConvertUnit
    2. 🔗 anthropics/claude-code v2.1.235 release

      What's changed

      • Added an optional spellcheck setting that underlines misspelled words in the prompt input as you type, using your installed aspell, hunspell, or ispell
      • Fixed whole-prompt-cache invalidation when a language server disconnected or reconnected mid-session
      • Fixed nested markdown list items misaligning at depth 3+ and added a hanging indent to wrapped list items in the terminal UI
      • Fixed prompt input highlights (slash commands, keywords, mentions) appearing shifted by one or more characters in some multi-line prompts
      • Fixed Shift+Tab inside the permission prompt's comment field approving the edit and granting session-wide edit permission instead of closing the field
      • Fixed the Agent tool advertising a general-purpose default in sessions where that agent is unavailable: an omitted subagent_type there now gets a clear error listing the available agents
      • Fixed notebook cell delete/replace approval dialogs silently omitting the existing cell content when the notebook or cell could not be read; the dialog now says why
      • Fixed slash commands run while Claude is responding showing HTML entities instead of the actual characters
      • Fixed the prompt footer not showing the "Update installed" restart notice after a background auto-update
      • Fixed the expanded task list (ctrl+t) always starting collapsed when resuming or relaunching into a session that still has open tasks
      • Improved memory and CPU usage while cloud sessions such as /ultrareview or /autofix-pr run in the background — their event streams are no longer re-scanned and re-rendered on every update
      • Improved permission dialogs: display text and "don't ask again" options now always match what a grant would cover, and "don't ask again" is withheld when contents cannot be fully displayed
      • Improved the embedded grep in native macOS/Linux builds: pathological patterns now fail fast instead of exhausting memory, and -m N with -A/-C prints correct context
      • Improved the context-limit error to say when auto-compact is off and point to /config to re-enable it
      • Vim mode: NORMAL mode and cursor position are now preserved when toggling the detailed transcript (ctrl+o) or closing a panel
      • Dialogs: arrow keys and Enter pressed in quick succession now select the option you navigated to instead of the previously highlighted one
      • SendMessage now refuses messages too large for cross-session delivery up front instead of silently dropping them
      • Remote Control: claude rc now applies the same enterprise-gateway availability check as interactive startup
      • [VSCode] Fixed focus jumping between open Claude tabs on its own when a window with several Claude panels is restored or reloaded
    3. 🔗 r/LocalLLaMA Alibaba's RISC-V CPU, XuanTie C950, Runs Qwen-3.8 27B at 30 tps rss

      Alibaba's RISC-V CPU, XuanTie C950, Runs Qwen-3.8 27B at 30 tps | Who needs GPUs? submitted by /u/DeltaSqueezer
      [link] [comments]
      ---|---

    4. 🔗 uswds/uswds USWDS 3.14.0 release

      What's new in USWDS 3.14.0

      Features

      Package | A11y | Breaking | Markup change | Description
      ---|---|---|---|---
      usa-accordion, uswds-core | Yes | Yes | Yes | Added a left-aligned expand/collapse icon option. A new $theme-accordion-icon-position setting (default: "start") lets teams set icon placement globally. The .usa-accordion--icon-start and .usa-accordion--icon-end modifier classes support per-instance placement. This improves discoverability for users viewing content at high magnification or zoom levels. Thanks @HopeTurnerUSCIS, @rosamundtgov, and @jeana-adhoc! (#6789)

      ✏️ Teams should verify layout at common zoom levels for the new default accordion behavior.
      usa-breadcrumb | Yes | Yes | - | Breadcrumbs now wrap by default. The previous truncation behavior is now opt-in via the new usa-breadcrumb--truncate modifier class. Thanks @AKnassa! (#6722)

      ✏️ Teams should confirm breadcrumbs display as expected and add usa- breadcrumb--truncate if they want the old truncation behavior.
      usa-range | Yes | - | Yes | Added a visible hint to the range slider. A new usa-hint element with the text "Move the slider to change the value" is added above the slider so sighted users receive the same guidance that screen reader users already had. Thanks @ravitejapioneerblaze-code! (#6673, #6811)

      ✏️ Teams should pull in the updated markup.
      usa-range | Yes | - | - | Improved range slider border visibility. The border is now 2px and uses the base-darker theme token. A focus ring is also added to the slider input. Thanks @manichandra! (#6659)
      usa-date-picker | Yes | - | Yes | Addedaria-current="date" to today's date button. Assistive technologies can now programmatically identify the current date in the calendar widget. Thanks @daresTheDevil! (#6593)
      usa-file-input | - | - | - | Error border now uses theerror-dark token. The file input error state previously used secondary-dark, which could show the wrong color in projects with distinct secondary and error palettes. Thanks @manichandra! (#6669)

      Bug fixes

      Package | A11y | Breaking | Markup change | Description
      ---|---|---|---|---
      usa-modal | Yes | - | - | Closing a modal now always restores screen reader access to the page. If the element that opened the modal had left the document by the time the modal closed, page content kept aria-hidden="true" and stayed invisible to assistive technology until reload. Thanks @vssinghh! (#6786)
      usa-modal, uswds-core | Yes | - | Yes | Fixed modal content being read twice by screen readers. The default focus target has changed — on open, focus now moves to the first enabled button in the modal footer, or the first enabled button in the modal if no footer button is present. FocusTrap no longer uses autoFocus. (#6703)

      ✏️ Teams should verify modal focus lands where expected after this update.
      usa-memorable-date | Yes | - | Yes | Added per-field hints to the memorable date component. The component previously used a single shared hint referenced by all three fields via aria-describedby, causing screen readers to repeat the full instruction block on every field focus. Each field now has its own targeted hint. The visible group hint remains for sighted users with aria-hidden="true". (#6725)

      ✏️ Teams should update to the new per-field hint markup. Teams supporting other languages should update their hint strings.
      usa-file-input | Yes | Yes | Yes | Removed the drag instruction on mobile and coarse-pointer devices. The file input previously showed "Drag file here or choose from folder" on all devices, including mobile where drag-and-drop isn't a practical interaction. Coarse-pointer devices now show and announce "Choose from folder" only. Thanks @manichandra! (#6660)

      ✏️ Teams should check for layout changes and update any accessibility tests that assert the old exact instruction text.
      usa-character-count | Yes | - | - | Deferredaria-live to prevent iOS VoiceOver from announcing character count on page load. Live updates continue to fire normally during typing. Thanks @daresTheDevil! (#6595)
      usa-character-count | - | - | - | Fixed the label selector so it correctly finds the associated label. Thanks @ealexhaywood! (#6385)
      usa-footer | - | Yes | - | Restricteddata-tag to heading elements. The footer's big link list was vulnerable to XSS through unsanitized data-tag values. Non-heading elements now gracefully fall back to h4. Thanks @IHIutch! (#6674)

      ✏️ Teams should review any footer data-tag values that aren't heading elements.
      usa-banner, uswds-core | - | - | - | Fixed toggle so it resolvesaria-controls from the component's root node. This allows banner toggles to work correctly when used inside a shadow root. Thanks @arpitjain099! (#6714)
      uswds-core | - | - | - | Fixed the language selector Escape key handler. Thanks @arpitjain099! (#6713)
      uswds-core | - | - | - | Guarded the keymap against non-keyboard events from datalist selections. This prevents a console error when a user selects an option from a datalist. Thanks @vijaygovindaraja! (#6594)
      usa-table | - | - | - | Restored the row header border in borderless tables. Row-scoped body header cells now keep their top border so row headers don't appear visually disconnected. Thanks @manichandra! (#6661)
      usa-table | - | - | - | Fixed.usa-sr-only table caption causing heading-row border collapse. Thanks @IHIutch! (#6633)
      usa-time-picker | - | - | - | Added missing combobox style dependency to the time picker package. Thanks @IHIutch! (#6634)
      usa-input, usa-textarea, usa-range, usa-combo-box, usa-input-prefix-suffix, usa-select | - | - | - | Setbox-sizing: border-box on the %block-input-styles mixin to prevent overflow. Components in host environments that reset global box sizing no longer overflow their containers. Thanks @VenkateshAddala! (#6736)

      ✏️ Teams should verify these elements display correctly in projects with custom global box-sizing resets (e.g. those who set $theme-global-border- box-sizing: false).
      usa-in-page-navigation | - | - | - | Standardized the component's enhancement guard to usedata-enhanced in line with other USWDS components. (#6688)
      uswds-core | - | - | - | Fixed ink assignment referencing the wrong variable. Custom ink colors in a project's theme were referencing the wrong system token. Thanks @nektro! (#6651)

      ✏️ Teams should verify that custom ink colors in their theme render as expected.

      Guidance changes

      Alert

      The alert component page now recommends that alert headings start with the alert type to improve clarity and urgency of the message, both for accessibility and general usability, as well as adding extended guidance to help teams use the correct alert type. Thanks @jeana-adhoc and @rosamundtgov! (#3288)

      Markup changes

      Memorable date

      The memorable date component's three fields now each have their own aria- describedby hint instead of sharing a single
      group hint. Teams who've copied the memorable date markup should update to the per-field pattern:

       <fieldset class="usa-fieldset">
         <legend class="usa-legend">Date of Birth&lt;/legend&gt;
      -  <span class="usa-hint" id="mdHint">For example: January 19 2000&lt;/span&gt;
      +  <span class="usa-hint" aria-hidden="true" id="memorable-date-hint">
      +    Select a month. Enter 1 or 2 digits for the day and 4 digits for the year.
      +  &lt;/span&gt;
         <div class="usa-memorable-date">
           <div class="usa-form-group usa-form-group--month usa-form-group--select">
             <label class="usa-label" for="date_of_birth_month">Month&lt;/label&gt;
      -      <select class="usa-select" id="date_of_birth_month" name="date_of_birth_month" aria-describedby="mdHint">
      +      <span class="usa-hint usa-sr-only" id="memorable-date-month-hint">Select a month from the dropdown.&lt;/span&gt;
      +      <select class="usa-select" id="memorable-date-month" name="memorable-date-month" aria-describedby="memorable-date-month-hint">
               ...
             &lt;/select&gt;
           &lt;/div&gt;
           <div class="usa-form-group usa-form-group--day">
             <label class="usa-label" for="date_of_birth_day">Day&lt;/label&gt;
      -      <input class="usa-input" aria-describedby="mdHint" id="date_of_birth_day" name="date_of_birth_day" ... />
      +      <span class="usa-hint usa-sr-only" id="memorable-date-day-hint">Enter 1 or 2 digits for the day.&lt;/span&gt;
      +      <input class="usa-input" aria-describedby="memorable-date-day-hint" id="memorable-date-day" name="memorable-date-day" ... />
           &lt;/div&gt;
           <div class="usa-form-group usa-form-group--year">
             <label class="usa-label" for="date_of_birth_year">Year&lt;/label&gt;
      -      <input class="usa-input" aria-describedby="mdHint" id="date_of_birth_year" name="date_of_birth_year" ... />
      +      <span class="usa-hint usa-sr-only" id="memorable-date-year-hint">Enter 4 digits for the year.&lt;/span&gt;
      +      <input class="usa-input" aria-describedby="memorable-date-year-hint" id="memorable-date-year" name="memorable-date-year" ... />
           &lt;/div&gt;
         &lt;/div&gt;
       &lt;/fieldset&gt;
      

      File input

      The file input no longer renders the drag instruction on coarse-pointer or mobile devices. On those devices, the
      instruction now reads "Choose from folder" instead of "Drag file here or choose from folder". For fine-pointer devices,
      the text remains "Drag file here or choose from folder." Teams with tests that assert the exact instruction text should
      update those tests.

      - Drag file here or choose from folder
      + Choose from folder
      

      Accordion icon alignment

      The default alignment for the accordion toggle icon switches to the left for better accessibility for users who zoom or
      use screen magnification. Teams can now use the usa-accordion--icon-start or usa-accordion--icon-end modifier to
      left-align or right-align the expand/collapse icon at the instance-level respectively. Icon position can be set globally
      with the $theme-accordion-icon-position Sass setting. The default behavior ("start" / left-aligned) is changed from
      v3.13.0, and the historical behavior can be preserved with $theme-accordion- icon-position: "end".

      - <div class="usa-accordion">
      + <div class="usa-accordion usa-accordion--icon-start">
      

      or

      - <div class="usa-accordion">
      + <div class="usa-accordion usa-accordion--icon-end">
      

      Or set globally in your theme:

      + $theme-accordion-icon-position: "start";
      

      or

      + $theme-accordion-icon-position: "end";
      

      Dependencies and security

      Dependency updates

      Dependency name | Previous version | New version
      ---|---|---
      lit | 3.2.1 | 3.3.3
      receptor | 1.0.0 | --

      Note: receptor has been removed as a dependency. Its functionality has been reimplemented in first-party code.
      Thanks @aduth! (#6489)

      Dev dependency updates

      Dependency name | Previous version | New version
      ---|---|---
      @babel/core | 7.26.8 | 7.29.7
      @babel/preset-env | 7.26.8 | 7.29.7
      @chanzuckerberg/axe-storybook-testing | 6.3.1 | --
      @material-design-icons/svg | 0.14.13 | 0.14.15
      @rollup/plugin-commonjs | 28.0.3 | 29.0.3
      @spiriit/vite-plugin-svg-spritemap | 4.0.0 | 6.0.0
      @storybook/addon-a11y | 6.5.16 | 9.1.20
      @storybook/addon-essentials | 6.5.16 | --
      @storybook/addon-links | 6.5.16 | --
      @storybook/builder-webpack5 | 6.5.16 | --
      @storybook/html | 6.5.16 | --
      @storybook/html-vite | -- | 9.1.20
      @storybook/manager-webpack5 | 6.5.16 | --
      @storybook/test-runner | -- | 0.23.0
      @types/node | 20.14.10 | 24.13.3
      @uswds/compile | -- | 1.3.2
      autoprefixer | 10.4.20 | 10.5.0
      axe-core | 4.10.2 | --
      axe-playwright | -- | 2.2.2
      concurrently | -- | 10.0.3
      css-loader | 6.8.1 | --
      del | 6.0.0 | 8.0.1
      esbuild | -- | 0.28.1
      eslint | 8.56.0 | 10.8.0
      eslint-config-airbnb-base | 15.0.0 | --
      eslint-config-prettier | 9.1.0 | 10.1.8
      eslint-plugin-airbnb-base | 0.0.1-security | --
      eslint-plugin-import | 2.31.0 | --
      eslint-plugin-import-x | -- | 4.17.1
      eslint-plugin-lit | 2.0.0 | 2.3.1
      eslint-plugin-no-unsanitized | 4.1.2 | --
      file-loader | 6.2.0 | --
      globals | -- | 17.8.0
      gulp | 4.0.2 | 5.0.1
      gulp-mocha | 9.0.0 | 10.0.1
      gulp-postcss | 9.0.1 | 10.0.0
      gulp-rename | 2.0.0 | 2.1.0
      gulp-sass | 6.0.0 | 6.0.1
      html-webpack-plugin | 5.6.3 | 5.6.8
      http-server | -- | 14.1.1
      magic-string | -- | 0.30.21
      merge-stream | 2.0.0 | --
      mocha | 10.8.2 | 11.8.0
      postcss | 8.5.2 | 8.5.25
      postcss-discard-comments | 6.0.2 | 8.0.2
      postcss-import | 15.1.0 | --
      postcss-loader | 7.3.3 | --
      postcss-preset-env | 9.6.0 | --
      prettier | 3.4.2 | 3.9.6
      react-dom | 17.0.2 | --
      resolve-url-loader | 5.0.0 | --
      sass-embedded | 1.83.4 | 1.100.0
      sass-loader | 16.0.4 | --
      sass-true | 6.0.1 | 10.1.0
      sinon | 12.0.1 | 22.1.0
      snyk | 1.1295.3 | 1.1306.2
      storybook | -- | 9.1.20
      style-loader | 3.3.3 | --
      svgo | 3.3.2 | 4.0.2
      twig | -- | 3.0.0
      twigjs-loader | 1.0.3 | --
      vite | 6.2.2 | 6.4.3
      vite-plugin-svg-sprite | 0.6.2 | --
      wait-on | -- | 9.1.0
      webpack | 5.98.0 | 5.109.2
      webpack-cli | 5.1.4 | 7.2.2

      0 vulnerabilities in regular dependencies (dependencies for USWDS projects installed with npm install @uswds/uswds)

      19 vulnerabilities (11 moderate, 3 high) in devDependencies (development dependencies)

      SHA-256 for release

      da91c65e6fc736fa397f0daf6ca2c2c95711506d85c7ad2537a4570725401e1c

      Additional contributions

    5. 🔗 exe.dev Have an Agent Babysit Your Deployments rss

      Deployments are scary. That’s the moment you break things.

      Not-deployments are even scarier. Waiting just makes the next deployment bigger.

      As the saying goes: “If it hurts, do it more.”

      The obvious, correct answer is CD. But then you’re off building canaries and waves and automated detection systems as gates. Canaries and waves are easy. Automated detection systems are hard. There are an indefinite number of things that can go wrong, and missing one of them takes you down. The asymmetry there is exactly the same asymmetry that makes deployments scary in the first place.

      The historical answer was: It’s a lot of engineering effort and a lot of pain. And so CD gets delayed, and humans babysit deployments, and deployments happen infrequently, and the cycle of inefficient misery and fear continues.

      This has exactly the right shape for an agent instead of code: Lots of rich data, a very long tail of possible states, relatively few runs (a handful a day, not 100qps).

      And we now have intelligence on tap. Let’s use it!

      At exe, Athena oversees our deployments. (All our bots have names, but that’s just so it’s easy to talk about them. They’re programs, not people.)

      Athena sits in a system called “exe-ops” which is our Deployment Command Center. We started with shell scripts, but then built a UI. Traditionally, you do a migration to something like Spinnaker. Instead, we’re building up from shell scripts into the exact shape we want. Athena is part of that story.

      The bot has read access to git and metrics and logs. It decides at each stage: should we proceed? Which machines should be in the next wave? It can escalate to a human and it can pause a deployment—or refuse to start one, if it deems it unwise. It communicates by sending us Slack messages.

      It’s great! It is diligent and thorough. It reads the diffs, analyzes the logs, checks for unforeseen issues, self-heals around weird problems, and reports on how to make future runs smoother.

      I could probably oversee deployments better than Athena. But the important question is not “in theory, could I do a better job?” but “in reality, will I do a better job?” We’re all busy. Athena does a much, much better job than I actually would.

      Athena lets me focus my attention elsewhere, until something happens that’s worth my intervention. And by deploying more often, those interventions are rarer and smaller.

    6. 🔗 r/LocalLLaMA Memory prices climb 500% in 12 months, up to 10x the lowest ever tracked prices - 128GB of DDR5 now $3,399 rss
    7. 🔗 HexRaysSA/plugin-repository commits sync repo: +2 releases rss
      sync repo: +2 releases
      
      ## New releases
      - [diaphora](https://github.com/joxeankoret/diaphora): 3.4.1
      - [eject_idb](https://github.com/allthingsida/eject_idb): 0.0.4
      
    8. 🔗 3Blue1Brown (YouTube) The jumping pegs puzzle rss

      Part of a series of monthly puzzles with MoMath.

    9. 🔗 @binaryninja@infosec.exchange What do Binary Ninja workflows do for you? A lot! Check out what mastodon

      What do Binary Ninja workflows do for you? A lot! Check out what @mei managed to pull off using them to clean up conditional jump threading:

      https://codeberg.org/mei-b/bn-analysis- improvements

    10. 🔗 r/LocalLLaMA Running DeepSeek V4 Flash Q4_K_XL at ~100 tok/s prompt processing on 4× RTX 3060 12GB rss

      Running DeepSeek V4 Flash Q4_K_XL at ~100 tok/s prompt processing on 4× RTX 3060 12GB | I managed to run the 143–144 GiB DeepSeek-V4-Flash-0731 UD-Q4_K_XL GGUF on four RTX 3060 12GB cards while keeping a 360k–376k context window. Hardware: CPU: Intel Core i9-10920X, 12C/24T RAM: 128 GB DDR4-3200, quad-channel GPU: 4× NVIDIA RTX 3060 12GB Total VRAM: 48 GB Storage: NVMe SSD Engine: llama.cpp, build b10181 Model: unsloth/DeepSeek-V4-Flash-0731-GGUF Quant: UD-Q4_K_XL, approximately 144 GiB KV cache: Q8_0 The best high-speed configuration so far: llama-server \ -m DeepSeek-V4-Flash-0731-UD-Q4_K_XL-00001-of-00005.gguf \ -c 368640 \ -ncmoe 34 \ -ts 100,1,1,1 \ -ot 'blk.(3[4-6]).ffn_.exps=CUDA1,blk.(3[7-9]).ffn.exps=CUDA2,blk.(4[0-2]).ffn.*_exps=CUDA3' \ -ctk q8_0 \ -ctv q8_0 \ -b 2048 \ -ub 2048 \ -np 1 \ -lm none \ --threads 20 \ --flash-attn on Measured with a roughly 20.5k-token prompt: Configured context: 368,640 tokens Prompt processing: 99.4 tok/s Text generation: 10.1 tok/s Minimum free VRAM under load: GPU0: 671 MiB GPU1: 842 MiB GPU2: 1395 MiB GPU3: 1395 MiB Model load time: approximately 198 seconds Other measured context/safety options: Context Prefill Decode Minimum free VRAM 376832 99.5 t/s 10.4 t/s 611 MiB 368640 99.4 t/s 10.1 t/s 671 MiB 360448 99.4 t/s 10.1 t/s 735 MiB The interesting part is the GPU layout. -ncmoe 34 keeps the experts from blocks 0–33 in system RAM. The remaining nine expert layers are explicitly distributed across GPUs 1–3, three layers per GPU. The extreme -ts 100,1,1,1 split does not distribute those explicitly assigned expert weights. Instead, it pushes most non-expert tensors—attention, KV-related allocations, etc.—onto GPU0. That leaves enough space on GPUs 1–3 for the large expert layers. This was much better than trying to calculate the layout analytically. With -ncmoe and explicit -ot overrides, tensor placement is discrete and somewhat unintuitive, so I measured every candidate. Microbatch size was the biggest performance lever: -ub 1024: approximately 63.4 tok/s prompt processing -ub 2048: approximately 99.4 tok/s prompt processing Decode remained almost unchanged at approximately 10.1–10.5 tok/s. At the full 393,216-token context, -ub 2048 also worked, but GPU0 had only 493 MiB free under load. Reducing the configured context to 368,640 restored a 671 MiB margin without reducing prompt-processing speed. For comparison, the safer -ub 1024 configuration can run with a configured context of 524,288 and still showed about 1032 MiB free on the tightest GPU, but prompt processing drops to approximately 63.4 tok/s. A few additional findings: Q8_0 KV is the default choice. F16 KV at c=393216 left only 587 MiB free. -ncmoe 33 caused a CUDA allocation failure. Memory mapping was disabled with -lm none. -np 1 is important; multiple slots multiply KV-cache requirements. The model is mostly in system RAM, so quad-channel memory bandwidth matters heavily. Even so, getting approximately 100 tok/s prompt ingestion and 10 tok/s generation from a 144 GiB MoE model on four consumer 12GB GPUs is much better than I expected. The configuration has been tested under real prompt load. The entire 368k context window has not yet been filled end-to-end, so the number above is the configured capacity, not a claim that I already completed a 368k-token generation test. Generated by ChatGPT 😂. submitted by /u/syscomua
      [link] [comments]
      ---|---

    11. 🔗 r/LocalLLaMA Qwen dev says not to wait for 35B-A3B rss

      Qwen dev says not to wait for 35B-A3B | What does this mean? Is there something else coming? Maybe 122B? Or no models? submitted by /u/Mean-Ad1493
      [link] [comments]
      ---|---

    12. 🔗 Ampcode News Education Discount rss

      Students and teachers can now subscribe to Amp for $10/month, half the usual price.

      What do you get for $10? Quite a lot!

      You get to use the best frontier agent. You get orbs, our remote machines that let you run agents from anywhere without supervision.

      You get code hosting for unlimited public/private repositories.

      And you get to use great models, through linking your ChatGPT sub for GPT-5.6 or 𝕏 Premium+/SuperGrok subscription for Grok 4.6. Plus $10 in credits each month for use on any other model.

      Get it at ampcode.com/edu.

    13. 🔗 Ampcode News Talk to Puck rss

      You can now talk with Puck in realtime:

      Realtime chat with Puck is powered by gpt-realtime-2.1. It delegates work to the Puck agent you already know, powered by GPT-5.6 Sol. Once the agent responds, gpt-realtime-2.1 summarizes the answer out loud while the full response appears in the thread.

      You can now have a proper back-and-forth conversation with Puck without waiting for text responses to stream in. Just talk. That makes it easier to coordinate parallel work, get progress updates, send follow-up instructions, or talk through big ideas without typing, whether you are sitting in front of your computer or on the go.

      Here are some conversation starters we have used:

      • "Check my active threads and tell me which ones need input."
      • "Review the messages in the #issues Slack channel where I am tagged. Read them to me one by one so we can talk through how to fix each one."
      • "Start an agent to fix the CI failure, tell me what caused it, and keep me updated on the fix."
      • "The executor lease reconciliation pipeline broke. Tell me about the changes made to it yesterday."
  4. August 17, 2026
    1. 🔗 IDA Plugin Updates IDA Plugin Updates on 2026-08-17 rss

      IDA Plugin Updates on 2026-08-17

      New Releases:

      Activity:

      • capa
        • b3817357: build(deps-dev): bump pyinstaller from 6.21.0 to 6.22.0 (#3157)
        • 84e36402: Sync capa-testfiles submodule
        • 1218e3b7: build(deps): bump setuptools from 83.0.0 to 84.0.0 (#3156)
      • chernobog
        • d272b5df: fix: stub HEXDSP in standalone tests to prevent crash on AST destruction
        • 79f1488d: ci: separate cache restore and save steps
        • 34c4f543: feat: expose full capability surface through IDC interpreter
      • disrobe
        • a5aa5b8e: recovery: expand appimage, flutter, php, jvm, and javascript
        • 65755cfa: recovery: recover python try loops, go control edges, and aarch64 ari…
        • d6d96518: javascript: recover system register parameter names
        • 4c01ca15: jvm: recover kotlin nested finally copies
        • 1471a1c1: query: expose canonical instruction effect rows
        • 92e5a255: javascript: recover rollup iife parameter names
        • 33e85455: recovery: expand erofs, go and php with bounded pdb argument lists
        • 8fa80485: recovery: expand erofs, nativeaot, flutter, python, php, wasm and d
        • a7a9fec7: recovery: expand firmware, python, javascript and witness output
        • 0e7c8980: recovery: add firmware, crc32 and report redaction
        • 81d68e0d: recovery: expand native, dalvik, php, as3 and wasm output
      • ffxiv_bossmod
      • ida-pro-mcp
      • plugin-ida
        • 0fe1e8c2: chore(deps): Bump step-security/harden-runner from 2.20.1 to 2.21.0 (…
      • project
        • fd1bb70b: updated webapppapebapp web notifications
      • twdll
        • 4a97c0e8: chore: update gitignore
        • 1837efd4: feat: add encyclopedia url func
        • 1dafd620: refactor: new database api
        • 02f96bbd: tests: improve set party test
        • e5abe584: refactor: structs refactor
        • 0a195e6b: chore: add manual test template
        • 3c7d1e70: feat: add SetPrimaryParty & SetParty for char
    2. 🔗 PrimeIntellect-ai/prime-agent v0.7.3 release
      • Fixed assistant rendering when provider payloads contain null or sparse content blocks.
      • Added authenticated host-request contracts with per-call request IDs, generation fencing, cancellation signals, and currentness checks.
      • Fixed root daemon shutdown retaining cleanup ownership while kill events are in flight.
      • Changed RLM family discovery to use a daemon-owned append-only spawn ledger with per-child display metadata instead of reconstructing topology from session files.
      • Fixed long-running macOS supervisors losing ownership when system cleanup removed authority records from $TMPDIR.
      • Fixed deleted RLM children leaking kernel snapshots while retaining their readable transcript tombstones.
      • Changed Agents View subagent rows to show stable name · model/effort · summary metadata.
      • Changed the default Cerebras model to the available gpt-oss-120b route and aligned cross-provider handoff fixtures with the generated catalog.
      • Fixed the agent going silent after an automatic context compaction interrupted unfinished work: the tool loop now resumes when a threshold compaction fails or is skipped, and active goals keep continuing after a successful mid-goal threshold compaction.
      • Changed the agents view splash hint from "type to start" to "type to search sessions".
      • Added app.edits.expand (ctrl+j) to toggle edit diffs; diffs are now shown only by this toggle, and ctrl+o no longer affects them.
      • Changed edit rendering so the ╰─ <path> +N -M summary line is always visible and ctrl+j toggles the diff inline beneath it, indented to the summary text.
      • Fixed fullscreen wheel scrolling in Ghostty while retaining application link clicks; set terminal.fullscreenMouse to false to use native Cmd-click instead.
      • Changed the agents view to sort idle and inactive sessions by last message time, newest first, while keeping running agents in stable creation order.
      • Fixed openai-codex models being invisible to rlm subagents and find_models because model discovery reported Prime Agent's own version as the Codex client version (#1375 by @bilelrais).
      • Added a working hint that recommends sharing traces with Prime Intellect to help train open-source LLMs.
      • Restored bare prime-agent --resume opening the agents view and the /resume [id|path] slash command; bare commands open the agents view and an argument resumes that session in place.
      • Fixed URLs not opening on click in fullscreen mode on terminals such as Ghostty; clicking a link in the transcript, dock, or overlays now opens it in the browser.
      • Fixed ctrl+p ("Toggle agent message expansion") only toggling received agent messages; it now expands and collapses sent agent messages together with received ones.
    3. 🔗 crosspoint-reader/crosspoint-reader v1.6.0rc release

      Summary

      This release is mostly bug fixes reported after 1.5.0. A handful of new features rode along too.

      New hardware support

      1.5.0 officially brought in support for our first ESP32S3-based reader — the Seeed reTerminal Sticky. With this release we now officially add support for the X4pro and M5Stack PaperMono devices. Both devices get full frontlight controls, swipe gestures in the reader, and the X4 Pro adds a capacitive Home key with configurable long-press settings.

      Want one? You can order them at crosspointreader.com/devices — that has our affiliate links, and buying from here helps support CrossPoint.

      Transparent sleep screens

      The sleep screen now supports transparent images. Add transparent PNGS and BMPs to your .sleep folder to see nice sleep screen overlays on top of your book pages!

      Reading Night Mode

      You can now toggle on Night Mode in the Reader settings to invert your display. This setting only effects your reader and will therefore flash white at the page refresh interval and revert to white outside of the reader and on your sleep cover.

      For dictionary users

      StarDict .syn synonym lookups are in, and HTML dictionary definitions now render through the EPUB engine instead of raw text — styled entries in dictionary lookups should display correctly now.

      The rest

      An Extra Wide line spacing option. You can now see the password while typing it into Wi-Fi, KOReader, and OPDS fields. Lists and tabs moved onto the new FUI framework.

      What's Changed

      New Contributors

      Full Changelog : v1.5.0...1.6.0rc

    4. 🔗 anthropics/claude-code v2.1.234 release

      What's changed

      • Added the optional CLAUDE_CODE_PROJECT_DIR_NAME environment variable: hosts that give each session its own config directory can choose a short name for the per-project transcript directory
      • Added the selection:clear keybinding action, so a key can be bound to clear an in-app text selection; also works in the agents view
      • Added a GitLab merge request badge to the footer and statusline: repos with a GitLab remote and an authenticated glab CLI show MR !N with draft/pending/green states
      • Claude Code now continues your session automatically when a claude.ai usage limit resets; turn it off in /config ("Continue automatically at usage limit")
      • Claude is now told to use your account email only to identify you, and not to send it to unrelated services unless you ask
      • Security: remote file reads, session restore, CLAUDE.md includes, workflow scripts and file uploads now reject Windows NT-namespace (\??\) paths, hardening the remaining pre-approval file accesses against the NTLM credential-leak vector
      • Fixed auto mode in very long sessions repeatedly re-checking and denying sandboxed commands' network access after the conversation had been compacted
      • Fixed session-scoped permission answers (including denies) being dropped when answering background subagent tool permission prompts
      • Fixed a crash when an API response on the non-streaming fallback path (typically via third-party gateways) contained a thinking block missing its thinking field or a text block missing its text field
      • Fixed markdown rendering becoming extremely slow for some messages containing unusual Unicode sequences
      • Fixed SendMessage rejecting a recipient copied from ListAgents when the session name is at the 200-character cap or emoji-heavy
      • Fixed repository detection mis-reading the host of git remotes with unusual userinfo, producing links and repo-specific behavior for the wrong host
      • Fixed MCP diagnostics printing resolved secrets: scope-conflict warnings now show the configured ${VAR} form, and connection-failure details show only the server origin
      • Fixed strictKnownMarketplaces allowlists accepting SCP-style git marketplace sources whose host differs from the one git would actually connect to
      • Fixed modal text such as the /login OAuth URL losing characters when copied in fullscreen
      • Fixed a --- horizontal rule in rendered markdown running into the line after it
      • Fixed consecutive shell commands splitting into multiple "Ran 1 shell command" rows when todo/task updates were interleaved between them
      • Fixed dialogs like /permissions opened while a ! shell command was running being dismissed when the command finished
      • Fixed a queued ! shell command being sent to the model as plain text after pressing up-arrow to edit the queued input
      • Fixed queued messages reappearing in the prompt history while still queued, Esc while selecting a queued message no longer interrupts the turn, and ! mode no longer sticks after a mid-turn submit
      • Fixed accepting the "Try the new fullscreen renderer?" prompt restarting the session without its permission mode (e.g. --dangerously-skip-permissions), tool allow/deny rules, model or effort flags
      • Fixed /tui dropping launch --allowed-tools/--disallowed-tools rules when it restarts; it now declines to switch, with the reason, when the session has restrictions a restart can't carry over
      • Fixed trust prompts omitting the repository-wide scope warning when the directory was first seen before the repository existed there
      • Fixed a case where an IDE diff tab closing during a permission re-prompt could answer the new prompt with the previous input
      • Fixed: files sent to the user during Remote Control sessions hosted by Claude Code Desktop or VS Code now upload, so they open on phone and web instead of showing an empty card
      • Fixed: after /login while CLAUDE_CODE_OAUTH_TOKEN is set, the stale-token reminder no longer leaks into Claude's automatically resumed turn — it now appears only to you
      • Fixed: permission previews now relay only to channel servers admitted by the inbound trust gate, and a server's explicit permission-capability opt-out is honored
      • Fixed: credential masking on relayed permission previews can no longer hide commands, paths, or destinations from the approver; oversized private-key blocks now redact under full-strength redaction
      • Fixed: provider API tokens that mask on permission previews now mask even when directly followed by shell delimiters
      • Fixed Claude Desktop inter-session messages being silently dropped by the recipient session when cross-session messaging read as disabled, which left the sender's query "thinking" for many minutes
      • Remote Control: signing this computer in to a different claude.ai account or organization now stops the running session within seconds and says why, instead of a misleading HTTP 404 hours later
      • Remote Control sessions started from Claude Code Desktop or VS Code now keep phones and claude.ai/code updated on the session's permission mode (and claude.ai/code on the model) as they change
      • Remote Control: effort picks made on a phone or on claude.ai/code now apply to terminal- and Desktop/VS Code-hosted sessions, and the session publishes its effort level to connected clients
      • SendMessage and ListAgents now say when your account's session list was too long to check completely, instead of treating unseen sessions as absent
      • Expired Anthropic profile credential now points you at /login when a claude.ai login would take precedence
      • Improved the transcript: your own prompts now render markdown (highlighted code blocks, inline code, lists) the same way replies do
      • Improved the "API returned an empty or malformed response" error to say what came back (content type, body kind, size, request ID) and why the original streaming request failed
      • Improved auto-generated session titles to read as short, specific names (e.g. "Login button bug") rather than sentences restating your request (e.g. "Fix the login button on mobile")
      • Reduced the context cost of loading the built-in claude-api skill from ~200k+ tokens to ~25k by loading reference docs on demand
      • /permissions can now be opened while Claude is working — rule changes apply to the rest of the current turn
      • /add-dir <path> can now be used while Claude is working; /add-dir, /autocompact, /theme, /help, /config and /advisor dialogs open mid-turn in the fullscreen TUI
      • /goal now clears itself with a notice when a turn dies on an unrecoverable error (e.g. revoked auth, an exhausted credit balance, or a context overflow) instead of staying armed
      • /goal: when background tasks keep a goal waiting for 30+ minutes, Claude now checks in on them instead of waiting indefinitely (set CLAUDE_CODE_GOAL_CHECKIN_MINUTES=0 to opt out)
      • claude setup-token now rejects unexpected extra arguments instead of silently ignoring them
      • Changed Esc in fullscreen mode to no longer clear a mouse text selection: it interrupts or dismisses as usual and the selection stays highlighted
      • Removed the redundant "Allowed by auto mode classifier" line that auto mode showed under every Agent tool call
      • Removed the "Default teammate model" setting from /config; agent-team teammates now use the leader's model unless the spawn names one
      • Dimmed the elapsed-time counter on the running tool header so it no longer competes with the bold counts
      • Background task notifications delivered between turns are now sent to the model inside <system-reminder> tags, matching mid-turn delivery
      • Mantle: skip the admin-pin availability probe at startup when a main-loop model is already picked
      • Windows: startup no longer stalls on repeated rename retries when ~/.claude.json is read-only
    5. 🔗 backnotprop/plannotator v0.27.4 release

      Follow @plannotator on X for updates

      Missed recent releases? Release | Highlights
      ---|---
      v0.27.3 | Folder watcher freeze fix on large repos, first SBOM-attested release pipeline
      v0.27.2 | Mobile plan and code review, Codex CLI 0.147 fix, folder annotate cold-start, configurable markdown extensions
      v0.27.1 | Open-in-editor launch fix, file headers respect Viewed/Git-add visibility toggles
      v0.27.0 | Call Flow analysis, --tailscale remote reviews, review panel remembers your view, Pi rebuild (breaking command rename), focus-mode shortcut
      v0.26.8 | Placed comment markers on HTML pages, shift-click multi-select, live app annotation
      v0.26.7 | Pinpoint targets any element on HTML pages, smarter hover labels, zero-scan hit testing
      v0.26.6 | Fixed empty environment variables in sandboxed sessions (Bun 1.3.14 builds)
      v0.26.5 | HTML pinpoint element annotations, durable annotate submissions, installer fallback for old git, vim HUD cursor fix
      v0.26.4 | Skill-menu hover jitter fix (same-day patch on v0.26.3)
      v0.26.3 | Skill references in comments with / or $, reachable remote session URLs, worktree switcher tooltips
      v0.26.2 | Single-file diff tabs render fully, no more silently dropped review files, light/dark theme pairs, palette-matched code blocks
      v0.26.1 | GitButler 0.22.0 compatibility via capability-probed JSON flags

      What's New in v0.27.4

      A Guided Review can now leave Plannotator. This release ships portable guide exports, share links on guides.show, and a guide CLI any agent can drive, alongside a favicon style switcher, jj support for Call Flow, GitLab artifact fixes in PR review, and a smoother call-flow Lens. Eighteen PRs, four from community contributors, two of them first-timers.

      Portable Guided Reviews and guides.show

      Guided Reviews used to live and die inside your review session. Now a guide has three ways out:

      Download it. Every guide gets a "Download portable guide" button that produces one HTML file containing the full guide and the diff it describes. It opens anywhere, renders exactly like the in-app guide with side-by-side diffs and per-section reviewed checkboxes, and needs no Plannotator install. The file stays small because it carries your content, not the renderer: the viewer loads from guides.show, pinned by filename and cryptographic checksum, so a tampered or wrong viewer never executes. Offline, the file degrades to a readable plain-text version of the guide.

      Share it. "Create share link" uploads the guide to guides.show and hands you a link anyone can open in a browser. Shares are end-to-end encrypted by default: the key lives in the URL fragment after the #, which browsers never send to the server, so guides.show stores bytes it cannot read. You also get a one-time delete token, and "Remove link" works from the same dialog for as long as that Plannotator remembers the share. An optional "Allow link previews" checkbox stores the guide unencrypted so chat apps can show its title; that is a choice, never the default. Setting PLANNOTATOR_SHARE=disabled turns all of this off.

      Author it from anywhere. The new plannotator guide subcommands (list, export, share, unshare) let any agent or script produce and publish a guide from a guide JSON and a patch, without a browser in the loop.

      Saved guides from v0.27.x load unchanged. The share service runs on Cloudflare with add-only, content-hashed viewer publishing and per-IP rate limiting on creation.

      Choose your favicon: Totman or the classic P

      The browser-tab icon is now a setting. Appearance settings offer two styles with visual previews: Totman, the current mascot, and Classic P, the original Plannotator mark restored byte-for-byte from the pre-mascot era. The server remembers your choice and serves it directly, so tabs show the right icon from the first paint without flashing the default. Hosts that embed the published UI packages are unaffected unless they opt in.

      Call Flow analysis on jj repositories

      Call Flow previously required a plain Git checkout. Reviews running on jj (Jujutsu) colocated repos now get the same changed-call-path analysis: the jj snapshot is resolved to the underlying Git objects and fed to the same CallDiff engine, with the same per-file Lens and dock views. Diff collection is untouched; this only extends where the analysis can run.

      GitLab PR artifacts fetch reliably and more safely

      Reviewing GitLab merge requests with uploaded artifacts (screenshots, logs, design files) got a hardening pass. Uploads now fetch through the authenticated API with a strict rewrite that only touches real upload URLs, falls back to the original web route when a self-hosted GitLab predates the API route, maps 401/403 responses to a clear "run glab auth login" hint, and no longer serves HTML or JavaScript content types through the artifact proxy. A regression test pins the invariant that credentials never follow a cross- origin redirect.

      The call-flow Lens stops fighting your scroll

      Community feedback within hours of trying Call Flow in Safari: the per-file Lens popover closed randomly mid-scroll and popped open for every badge that passed under the cursor. Three causes, three fixes: the Lens's internal scroll no longer chains to the page when momentum hits its edge (the chain moved the popup out from under a stationary pointer, which read as a random close and was worst under Safari rubber-banding); hover now has a 100ms intent delay so drive-by badges stay closed; and an in-flight page scroll holds any pending close until the scroll settles.

      Reported by Rustan (@acewhocares on X).

      Additional Changes

      • Touch selection survives the comment composer. On phones and tablets, dragging a multi-line range in a single-file diff no longer collapses the selection when the composer opens; the range you dragged is the range you comment on. #1333
      • Skill picker works with screen readers. The / and $ skill reference menu now exposes real listbox semantics with option roles and active-descendant tracking, so assistive tech announces what Enter will insert, closing #1233. #1316 by @ashish921998
      • Blog: an interactive UI for the grill-me skill. A new post on using /plannotator-last as the review surface for Matt Pocock's grill-me workflow, at plannotator.ai. #1321, #1322, #1323, #1332
      • Security page linked from the site footer. #1305

      Install / Update

      macOS / Linux:

      curl -fsSL https://plannotator.ai/install.sh | bash
      

      Windows:

      irm https://plannotator.ai/install.ps1 | iex
      

      Claude Code Plugin: Run /plugin in Claude Code, find plannotator , and click "Update now".

      OpenCode: Clear cache and restart:

      rm -rf ~/.bun/install/cache/@plannotator
      

      What's Changed

      • fix(comments): expose skill picker semantics to assistive tech by @ashish921998 in #1316
      • feat(review): jj support for Call Flow analysis by @graemefolk in #1312
      • blog: an interactive UI for the grill-me skill by @backnotprop in #1321
      • blog: grill-me post additions by @backnotprop in #1322
      • blog: repo link, image alt text, and larger blog type by @backnotprop in #1323
      • feat: Portable Guided Reviews, export, share links, agent-authored guides, guides.show by @backnotprop in #1324
      • guides-show: GitHub link in the landing page header by @backnotprop in #1327
      • guide-viewer: label agent harnesses in the generated-by line by @backnotprop in #1328
      • guide-viewer: readable on phones and tablets, desktop untouched by @backnotprop in #1329
      • seo: index live root blog pages by @backnotprop in #1332
      • guide: voice rules in the organizer prompt by @backnotprop in #1330
      • docs(marketing): link security page from footer by @backnotprop in #1305
      • fix(review): preserve dragged diff ranges on compact touch before commenting by @backnotprop in #1333
      • feat(ui): Totman/Classic P favicon style switcher by @FNDEVVE in #1325
      • fix(review): GitLab upload artifact fetching via authenticated API with hardened rewrite by @yuensunn in #1228
      • guides-show: example guide screenshot at the bottom of the landing page by @backnotprop in #1336
      • guides-show: example guide screenshot replaces the abstract figure, opens in a lightbox by @backnotprop in #1337
      • fix(review): stop the call-flow Lens closing mid-scroll and opening on drive-by hovers by @backnotprop in #1338

      New Contributors

      Contributors

      Four community authors shipped code in this release, two for the first time:

      • @FNDEVVE built the favicon style switcher in #1325, including restoring the classic P icon exactly as it shipped before the mascot era, and worked through a review round that added server-side icon serving so the choice applies without a flash. Their second contribution.
      • @graemefolk extended Call Flow analysis to jj repositories in #1312, their third contribution to Plannotator's jj support, which they have carried since the original provider landed.
      • @yuensunn fixed GitLab merge request artifacts in #1228, their first contribution, and stuck with it through a security-focused review round on the URL rewrite.
      • @ashish921998 made the skill reference menu real for screen reader users in #1316, their first contribution.
      • Rustan (@acewhocares on X) test-drove Call Flow in Safari and reported the Lens scroll behavior that #1338 fixes, hours after trying the feature.

      Full Changelog : v0.27.3...v0.27.4

    6. 🔗 r/LocalLLaMA Artificial Analysis' Qwen3.8-27B benchmarks put it neck and neck with DeepSeek V4 and GPT-5.6 Luna Max rss
    7. 🔗 @binaryninja@infosec.exchange Seven sidebars was a bit much. Sidekick 26.1 brings Indexes, Notebooks, Code mastodon

      Seven sidebars was a bit much. Sidekick 26.1 brings Indexes, Notebooks, Code Maps, and Repositories together in one Sidekick Resources sidebar! Search across all four from the same place, then open or pin whatever you need right in Binary Ninja. See what else is new in 26.1: https://sidekick.binary.ninja/blog/sidekick-26-1-a-proper-home-for- sidekick/#seven-sidebars-were-too- many

    8. 🔗 @malcat@infosec.exchange Did you know that [#Kesakode](https://infosec.exchange/tags/Kesakode) can use mastodon

      Did you know that #Kesakode can use fuzzy-matching for functions? While a bit slower, this kind of lookup helps against obfuscation (here: #vidar)

    9. 🔗 r/LocalLLaMA After pushing 1M+ tokens through Qwen 3.8 27B, here is my optimal llama.cpp config for 16GB VRAM (73k Context, Agentic Coding) rss

      After pushing 1M+ tokens through Qwen 3.8 27B, here is my optimal llama.cpp config for 16GB VRAM (73k Context, Agentic Coding) | Dando seguimiento a mi post anterior sobre cómo tengo montado mi servidor de presupuesto (Intel N100 + RTX 5060 Ti 16GB), varios me preguntaron por una mirada más profunda a mi configuración real de inferencia y al desempeño agentic en el mundo real. Como muchos de ustedes, estaba refrescando la página esperando descargar Qwen 3.8 27B apenas salió. Después de pasar todo el fin de semana estresándolo con flujos de trabajo de codificación agentic, logré correr un proyecto completo y grande casi todo de forma autónoma (más de 1M de tokens procesados en total , solo 3 prompts). Aquí va un resumen rápido de la configuración base antes de meternos en los detalles del config y del workflow.

      Specs y parámetros rápidos

      • Modelo: Qwen3.8-27B-UD-Q3_K_XL.gguf
      • Hardware: RTX 5060 Ti (16GB VRAM) + Intel N100 (4C/4T, 16GB RAM)
      • Ventana de contexto: 73,728 (73k de contexto) corriendo tranqui en 16GB de VRAM.
      • Cuantización de KV Cache: q4_1 para el contexto principal
      • Decodificación especulativa: MTP nativa activada (spec-type = draft-mtp, n-max = 2)
      • Sampling: temp = 0.65, top_p = 0.95, top_k = 20, min_p = 0.05

      El experimento: armar una API completa con 3 prompts

      En vez de correr benchmarks sintéticos, metí esta configuración por una cadena real de ingeniería de software: construyendo una REST API no oficial y un servidor MCP para un foro vBulletin heredado.

      1. Prompt 1 (Arquitectura del sitio y análisis): Pedí al modelo que mapee el sitio objetivo. Generó una especificación en Markdown impecable de ~1,500 líneas que cubría análisis estructural, nodos HTML rescatables, payloads JSON esperados, selección de stack, lógica de paginación, autenticación de sesión y endpoints de búsqueda—mucho más a fondo de lo que yo habría escrito a mano.
      2. Prompt 2 (Arquitectura de desarrollo): Usando la spec como única fuente de verdad, diseñó un plan de implementación modular de NestJS dividido en 9 fases de ejecución:
      3. Fase 1: Estructura inicial del proyecto
      4. Fase 2: Modelos de dominio
      5. Fase 3: Scraping core (HTTP + limitación de tasa + reintentos)
      6. Fase 4: Parsers de HTML (cheerio)
      7. Fase 5: Capa de caché
      8. Fase 6: Servicios de aplicación + REST API
      9. Fase 7: Autenticación (sesiones con cookies)
      10. Fase 8: Servidor MCP (entrega principal)
      11. Fase 9: Fortalecimiento, documentación y entrega
      12. Prompt 3 (Ejecución autónoma agentic): La prueba de verdad. Le pedí a OpenCode (usando Qwen 3.8 27B) que actuara estrictamente como orquestador, creando sub-agentes para cada fase de tareas. Corrió de forma autónoma por ~2 horas. Cuando se acercaron los límites de contexto, OpenCode resumió su estado y siguió construyendo. Escribió tests unitarios, aplicó linting y entregó código 100% funcional—solo necesitando un arreglo automatizado menor cuando le di un payload de HTML crudo con un caso extremo.

      El archivo de configuración llama.cpp

      Aquí está mi archivo exacto de configuración de enrutador --models-preset . Fíjate cómo fit = off se usa en el perfil de 27B junto con ctx-size = 73728 (73k) y q4_1 para cuantizar la KV cache, con el objetivo de maximizar la asignación de VRAM mientras se mantiene el rendimiento nativo de MTP. ```ini

      ==============================================================================

      LLAMA.CPP — CONFIGURACIÓN DE INFERENCIA (modo router / --models-preset)

      ==============================================================================

      Objetivo de hardware:

      GPU: 16 GB VRAM (RTX 5060 Ti)

      CPU: Intel N100, 4C/4T (Debian Headless)

      ------------------------------------------------------------------------------

      GLOBAL / LÍNEA BASE

      ------------------------------------------------------------------------------

      [*]

      --- HILOS DE CPU


      Reserva 1 core para SO/servicios durante el decode.

      Usa los 4 threads durante ráfagas de prefill del prompt.

      threads = 3 threads-batch = 4

      --- SERVIDOR / CONCURRENCIA


      Un solo slot; desactivado continuous batching para máximo rendimiento por

      usuario.

      parallel = 1 cont-batching = 0

      --- GPU / AJUSTE DE VRAM


      flash-attn = on fit = on

      Holgura de seguridad para el límite físico de VRAM (MiB).

      Ponlo bajo (128) porque el sistema es headless (100% VRAM disponible para

      inferencia).

      NOTA: Si usas caches KV draft de MTP, ojo con la asignación doble de VRAM.

      Sube a 128-256 si te topas con OOMs.

      fit-target = 128

      --- CONTEXTO & CACHÉ ------------------------------------------------------

      ctx-size = 65536 context-shift = 1

      Desactiva checkpoints de contexto (evita problemas de reprocesamiento en

      arquitecturas híbridas)

      ctx-checkpoints = 0

      RAM Prompt Cache (2 GiB)

      cache-ram = 2048

      --- KV CACHE GLOBAL


      cache-type-k = q5_1 cache-type-v = q5_1

      --- PREFILL / BATCHING


      batch-size = 2048 ubatch-size = 1024

      --- SAMPLING POR DEFECTO (Códigos / Precisión)


      temp = 0.5 top-p = 0.95 top-k = 20 min-p = 0.05 repeat-penalty = 1.0

      ------------------------------------------------------------------------------

      QWEN 3.8 27B — PERFIL DE RAZONAMIENTO & CODIFICACIÓN PESADA

      ------------------------------------------------------------------------------

      [qwen3.8-27b] model = /opt/llama- infrastructure/models/Qwen3.8-27B-UD-Q3_K_XL.gguf

      Desactiva "fit" para evitar que capas se carguen en la CPU por un error de

      cálculo automático

      fit = off ctx-size = 73728 context-shift = 1

      MTP nativa del modelo (Decodificación especulativa)

      spec-type = ngram-mod,draft-mtp spec-draft-n-max = 2

      Cuantización de KV (q4_1 nos permite meter contexto de 73k en 16GB de VRAM)

      cache-type-k = q4_1 cache-type-v = q4_1

      Parámetros de presupuesto de pensamiento / razonamiento

      chat-template-kwargs = {"preserve_thinking": true, "reasoning_effort":"medium"} reasoning-budget = 5000

      Batches más chicos para evitar picos de VRAM durante prefills masivos

      batch-size = 1024 ubatch-size = 512

      Ajustes oficiales / recomendados del sampler de cuantización

      temp = 0.65 top-p = 0.95 top-k = 15 min-p = 0.05 ``` submitted by /u/chiribe
      [link] [comments]
      ---|---

    10. 🔗 seanmonstar Tending my little plot of the Internet rss

      As many parts of the Internet continue to get worse, I figured it was time to improve my own little plot.

      I mean, I’m always tinkering, making small tweaks regularly. Did you know I keep /now up-to-date? But anyways, a couple changes here felt big enough to write about.

      My microblog is mine

      I’ve been outputting random microposts since… checks archive 2009, apparently. They were “status updates” back then. With Twitter dying, I started doing that sort of thing on Mastodon, and then on BlueSky. Wherever the people want to be, I suppose.

      At first they were just silly jokes. But eventually, besides announcements, they became ways to express raw (bad) ideas and get feedback. But something about that always bugged me: they were on someone else’s property, and linking to them (let alone finding them again) felt bad.

      So, I own my microblog now. They have their own place on this domain. They get a dedicated RSS feed. And they are included in the main feed (currently prefixed as “Micro” so you know). They get syndicated to those other networks automatically as threads.

      What makes them micro? I don’t constrain myself to just 250 characters or anything. They’re so far about 3 paragraphs. That’s about the size, I aim for, I guess. It let’s me publish thoughts without blocker energy telling me I need to polish it into an essay. It also allows me to output 1 or 2 a week. Or none.

      And I can link to them and build on them. Mine!

      Subscribe via email

      I have improved the subscribe via email option of this site.

      For a long time, “subscribe via email” was easy and nice, provided by Feedburner. When that service was shutdown, I looked for an alternative. Something that was both free and automatically just worked from an RSS feed.

      I’m sorry about that. I picked something horrible, a service I don’t want to provide any further attention. They inject gross click-baity ads inside the emails. I subscribe to myself, and after being repulsed at the last email, I had to fix it.

      I couldn’t find any other service that automatically works from RSS for free. I could pay for a service, but I don’t need to send that much email. And I’m doing this as a convenience, to let users read how they want, not as a business. It’s not a newsletter.1

      So I imported that list to Buttondown. I can copy-paste the markdown of blog posts manually. That’s fine, I don’t write so often to need it to be automated.

      But since it is manual, I can do more. I can also include a list of “recent microblog posts”, now that I own them.

      Anyways, back to continual tinkerage.2

      1. And there’s no way I could subject my readers to Substack or Medium or something. Those sites do not treat readers well. I automatically refuse to read any article on such a site. I assume that if you don’t care about my reading experience, I don’t care enough about your idea. Not sorry. 

      2. Other things I want to improve: a combined blog and micro archive. A tags page. A better chronological story for About. A set of “values” pages. 

    11. 🔗 r/LocalLLaMA Petition to add a rule for people to add their DAMN quant levels to their posts rss

      Every time I see a post about a newly released model, whether it be a comparison or shitting on it, I have to dig through the endless comments to see what quants they used and what their specs were.

      Its quite a common occurrence here in this sub to ask someone that's saying a model is underperforming, and when you ask what quantization they are running they say something like "oh im running q0.1bpw from nobodyknowswhothisguyis".

      Worst offender is with comparison posts. "Comparing the new Qwen3.8-27B to Qwen3.5-9B and the 9B model is better!" I wonder why?

      Sorry for bad england

      submitted by /u/Su1tz
      [link] [comments]

    12. 🔗 r/LocalLLaMA Stripe will reportedly acquire AI gateway startup OpenRouter for $7B+ rss

      another one ..

      submitted by /u/ab2377
      [link] [comments]

    13. 🔗 r/LocalLLaMA …and I’m not afraid of losing my social credits. rss

      …and I’m not afraid of losing my social credits. | submitted by /u/JLeonsarmiento
      [link] [comments]
      ---|---

    14. 🔗 Filip Filmar Bazel all the way down: how I build programmable hardware rss

      This is a description of how I build programmable hardware. Everything that goes into Cocoapuffs, my RISC-V system-on-chip on an Artix-7 FPGA: the RTL, the firmware, the simulations, the synthesis, the bitstream, and the programming of the board, comes out of a single bazel build, from a machine that has nothing installed on it but bazel. The build is hermetic, ephemeral, and reproducible, and it is the same build whether it runs on my laptop, on a virtual machine in the cloud, or in continuous integration.