- ↔
- →
- October 02, 2026
-
🔗 New Music Releases Imminence - Axis Mundi rss
Imminence - a new release is available:
- 2026-10-02: Axis Mundi (Album)
Amazon: Canada | Deutschland | France | United Kingdom | United States
Visit muspy for more information.
-
- October 01, 2026
-
🔗 earendil-works/pi v1.0.0 release
New Features
- Fullscreen by default — The TUI now runs fullscreen. Set
tuiModeto"regular"to keep the terminal's normal scrollback. See Terminal and display. - Leaner codemode — About 40% fewer prompt tokens, and errors that tell the model how to recover. See Codemode.
- Image generation in codemode — Scripts call
models.generateImages()with the session's credentials. See Generate images and Use image models. - Radius in
/login— Sign in with Radius and set up its MCP server in one step. See Radius. - Anthropic copy code login — Sign in when the browser runs on another machine. See Authenticate interactively.
- MCP OAuth hardening —
oauth.authServerMetadataUrl, RFC 9207isschecks, credentials per server, and step-up sign-in that keeps granted scopes. See Authenticate with OAuth. - Header-only quiet startup —
quietStartup: "header"keeps the version and key hints and hides the rest. See Terminal and display.
Added
- Added an
oauth.authServerMetadataUrlsetting for MCP servers that advertise a wrong OAuth authorization server or none. Pi uses the configured metadata document instead of discovery (#10172). - Added
quietStartup: "header", which keeps the startup header with version and key hints but hides the model scope line and loaded-resource listing. - Added
models.generateImages()to codemode scripts. It runs image models such as OpenRouter's with the session's credentials and returns base64 image blocks thatimage()attaches to the result; usage counts toward the session cost likemodels.classify(). Extensions can callctx.modelRegistry.generateImages(). See Use image models. - Added a copy code login method to Anthropic
/loginfor headless setups where the browser runs on another machine (#10194 by @lucasmeijer).
Changed
- Changed the default TUI mode to fullscreen. Set
tuiModeto"regular"or pass--tui-mode regularto keep the terminal's normal scrollback. /loginnow offers "Sign in with Radius" at the top level, as the last option, with its status. After a Radius sign-in,/loginoffers to configure the Radius MCP server in the globalmcp.jsonwith"auth": { "provider": "radius" }and reloads. Cancelling a login returns to the menu it was started from.- The provider docs page is renamed to Providers, its "Cloud Providers" section is now "Provider Specific Config", and it documents Radius first.
- MCP OAuth credentials are now stored per server name and URL, so MCP servers with the same URL can sign in with different accounts. Credentials stored by URL alone move to the first server that uses them (#10252).
- Codemode costs far fewer prompt tokens: with the default tools and codemode active, a GPT-5.6 request shrinks from about 5,300 to 3,300 tokens. The
codemodedescription lists the script globals in one line each and points to the new Codemode reference for themodelsAPI, which the model reads when it needs it. Declared tools say in one line how scripts call them and what the call resolves to, instead of repeating their full declaration, and the system prompt's codemode guidance and MCP server section are shorter. - Codemode errors now say how to recover: reading a tool or
modelsmember that does not exist names the close matches (tools.Bashsuggeststools.bash),models.classify()andmodels.generateImages()reject malformed arguments with the expected shape, an unknown model points tomodels.getAvailableOfType(), an oversizedstore()value explains what the store is for, and a script that generates images without showing them gets a note. Scripts that probed for a tool withtypeof tools.namemust use"name" in tools. /loginand/logoutnow label providers without credentials as "not configured" instead of "unconfigured".- OAuth browser pages now show the color Pi logo.
Fixed
- Fixed MCP OAuth sign-in accepting an authorization response whose
issparameter names another authorization server; the code is now rejected before it is exchanged (RFC 9207). - Fixed MCP OAuth sign-in failing with
Invalid scopewhen the token response contains"scope": "", and similar failures for other empty ornulloptional OAuth fields (#10266). - Fixed the sign-in URL printed by
/mcp loginnot being clickable when it wraps (#10186). - Fixed
--providerwithout--modelbeing silently ignored and running the default model from another provider; it now fails with an error (#10236). - Fixed MCP servers that ask for more scope (
insufficient_scope) requesting sign-in over and over. The new sign-in requested only the missing scopes, so the new token lost access the previous one had; it now keeps the granted scopes. - Fixed user messages in the transcript keeping two full-width copies of every rendered line; they keep one, with identical output.
- Fixed
/loginand/logoutlabeling every OAuth sign-in, including Radius, as a subscription; only subscription-backed providers say "subscription", other OAuth sign-ins say "account". - Fixed the startup header logo rendering with gaps in Apple Terminal; it now shows a colored "Pi" with the version instead.
- Fixed deferred MCP tools that
tool_searchloaded being dropped on resume and/reloadeven when their server reconnected before the next prompt, because the session restored its tools before the MCP servers reconnected. - Fixed the system theme making pastel palettes such as Catppuccin Frappe much more vivid; palette colors now keep their chroma (#10255, #10293 by @dgtlntv).
- Fixed slash command autocompletion not triggering when the input starts with whitespace (#10218 by @haoqixu).
- Fixed color bleeding past mouse selections and search highlights in fullscreen mode when a styled token ends at the highlight boundary (#10169).
- Fixed memory retained per rendered message in the transcript; a long assistant message keeps about a fifth of the heap it kept before.
- Fullscreen by default — The TUI now runs fullscreen. Set
-
🔗 anthropics/claude-code v2.1.287 release
What's changed
- Added Claude Mods: plugins may now modify deeper behavior
- Added You should know, a built-in mod where a side agent watches your back and flags things you or Claude might miss. Turn it on with
/plugin enable cc-plugin-you-should-know@builtin(for first-party sessions with telemetry on) - Added an
n:<text>filter to the agents view that matches session names and tasks; a filter now shows matches in collapsed sections and Enter opens the first match - Added
prompt_textto the OpenTelemetryuser_promptevent, a copy ofpromptfor backends that nest dotted keys; drop or mask it wherever you drop or maskprompt(#70763) - Added URL prompts from MCP servers on the 2025-11-25 protocol, for example to sign in. If a server no longer connects after this update, add "bareElicitationCapability": true to its MCP config entry
- Windows: Added a startup warning when denying the Bash tool also turns off the PowerShell tool, so Claude has no shell tool
- Self-hosted runner: Added a built-in
gh api(REST only) for sessions that use Anthropic-managed git on macOS and Linux machines where the GitHub CLI is not installed - Fixed fast mode staying off in remote sessions owned by an agent with no user account, even when the organization allows it
- Fixed Remote Control not receiving messages for minutes at a time when a reconnect request got no response; it now gives up after 30 seconds and retries
- Fixed hooks configured with
asyncRewakewaking Claude over and over with "found issues" notifications when the hook's script file is missing; the broken hook is now reported once - Fixed tool heartbeats not reaching SDK hosts while the model's response stream was stalled with no data arriving
- Fixed Bedrock and Vertex startup model checks ignoring an enforced
availableModelslist, which could collapse/modelto one Opus row - Fixed the Claude in Chrome browser picker showing a JSON parse error when Chrome could not be reached
- Fixed picking Fable in
/modelon a claude.ai login saving the current version's id, so your saved default now follows the newest Fable like Opus and Sonnet do - Fixed switching between Opus 5.5 and Sonnet 5.5 (
/model,opusplan) rewriting earlier MCP tool announcements, which could drop earlier extended thinking - Fixed Amazon Bedrock Guardrails blocks that arrive mid-response ending the turn with an API error instead of the guardrail's message when the reply began with thinking
- Fixed a dangerous
rm(such as one on/or the home directory) losing its always-ask safeguard when the same command also redirected output to a~or wildcard path - Fixed
claude -pand SDK sessions repeating a model fallback on every later message after the model was switched while a reply was running - Fixed a folder's CLAUDE.md being attached a second time after resuming a session or after a compaction
- Fixed background sessions that could not be reopened from
claude agentsafter the agent exited and removed the worktree the session was started in - Fixed
/advisorpairing checks: Sonnet 5.5 can now advise Opus 4.7 and 4.8, and advisors the API would refuse are flagged up front instead of being silently dropped - Fixed Bash permission prompts showing internal parser names such as "Contains simple_expansion" instead of a plain explanation
- Fixed a cause of fullscreen sessions on slow or busy machines exiting with "Claude Code exited after an unrecoverable interface error" while a scroll key was held in a long conversation
- Fixed organization per-tool permission ceilings being silently dropped for an MCP tool named
__proto__ - Fixed Claude being told to page large MCP results saved as JSON with Read's offset and limit, which cannot split one long line
- Fixed the commit attribution reminder being delivered inside a tool result after a compaction
- Fixed screen reader mode leaving the cursor away from the typed text in search boxes (such as /resume and /permissions) and sign-in code fields
- Fixed screen reader mode refusing Enter with nothing typed on /rewind's summarize options, whose added context is optional
- Fixed screen reader mode showing a "Tab to amend" hint on approval prompts, where Tab does nothing
- Fixed screen reader mode listing arrow keys that do nothing in /permissions and /mcp, and saying "Select with numbers" in empty menus or while a search box has the keys
- Fixed screen reader mode leaving out the changed lines in file edit approval prompts and other diffs
- Fixed screen reader mode sending the
claude --teleportprogress screen, and an MCP form field while it is being checked, to the screen reader again on every spinner frame - Fixed
CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETASnot removing the structured-output format from session-title and prompt-hook requests, which Bedrock-backed gateways reject - Fixed screen reader mode leaving out the top lines of a second approval prompt, a changed /config row or the rejected-plan line when the previous screen was taller than the terminal window
- Fixed
--include-partial-messagessending a cut-short reply'smessage_stoplate or never, so apps could show the reply as still in progress - Fixed
claude agentssometimes not showing the permission prompt a background session is waiting on - Fixed
/ultrareviewgiving advice about.git/info/attributeswhen the upload stops on a committed.gitattributesit cannot read, such as one saved as UTF-16 - Fixed
claude remote-controlfailing to register behind an HTTP proxy with a misleading "Check your organization permissions" error (#97352) - Fixed sandboxed Bash commands on Linux inheriting an open handle on the Claude Code executable
- Fixed the running-tool dot and three spinners still moving with the "Reduce motion" setting on, and /rewind's confirm screen updating its "ago" time while you type a note
- Fixed times in
claude agentschanging every second in screen reader mode; they now change at most every 10 seconds - Fixed a revoked claude.ai login showing a generic
API Error: 401instead of "OAuth token revoked"; in-pmode the error now starts with "Failed to authenticate" - Fixed
/ultrareviewupload refusals advising you to copy a variable named by a repository's settings file into your own user settings - Fixed
--output-format stream-jsonand the SDK not streaming the turns of acontext: forkskill run by typing/<skill>as the prompt, as they do for the Skill tool's fork - Fixed /feedback and /bug: the pre-filled GitHub issue no longer includes your recent error messages, and the confirmation screen now lists them as part of the report
- Fixed
claude plugin marketplace add --sparseandgit-subdirplugin installs failing with "transport 'http' not allowed" when the repository is served over plain http - Fixed cloud sessions sometimes losing the earlier conversation when the session restarted while it was being compacted
- Fixed a plugin reload that overlapped the startup
--plugin-urldownload corrupting the session's cached copy of the plugin archive - Fixed
/desktopquoting partial output when opening Claude Desktop timed out or printed too much output; the error now names the cause - Fixed an MCP connector tool call occasionally running twice, or the connector's calls failing until restart, when its server changed which MCP protocol version it supports
- Fixed SessionStart hooks from synced plugins not running in new cloud sessions
- Fixed the transcript's "N hooks ran" summary and the verbose debug log's matched-hooks count including Claude Code's internal callbacks, so one configured hook no longer shows as two
- Fixed files Claude sends from cloud and Remote Control sessions failing when the upload finished just after the 30-second timeout; it now waits 35 seconds
- Fixed repositories added mid-session in cloud and SDK sessions not loading their skills and plugins, and loading CLAUDE.md late, after Claude changed directory
- Fixed PNG, JPEG and WebP images over 8,000 pixels on a side failing to send from a remote session; Claude now sends a scaled-down copy
- Fixed messages sent from the Claude apps with 17 to 20 attached files delivering only the first 16
- Fixed headless sessions reporting an MCP server as needing authentication after one refused call, even though later calls succeed
- macOS: Fixed Remote Control sessions started with
claude remote-controlstopping mid-turn when the Mac went to idle sleep - Windows: Fixed interactive
claudehanging or crashing with "Raw mode is not supported" when its input is piped or redirected; it now says why and exits (use-pfor piped input) - Bedrock, Vertex, Mantle: Fixed model availability checks under
CLAUDE_CODE_SKIP_*_AUTHsending a differentAuthorizationheader than real requests whenANTHROPIC_CUSTOM_HEADERSrepeats it - Improved
/config: settings that cycle show ‹ › and step both ways with ←/→, narrow terminals stack each value under its label, and PgUp/PgDn page the list - Improved plugin marketplace errors to say in plain words why a marketplace was ignored or refused, and what to do
- Improved plugin listings to note when a plugin's dependencies were not installed, and updating a plugin now retries an install that did not finish
- Improved the Claude apps gateway's error when Amazon Bedrock rejects a model ID: developers now see which model is unavailable, and the gateway log names the ID that was sent
- Improved SDK sessions so a message sent with priority "now" no longer cancels a running web fetch or web search; it keeps loading in the background
- Improved
/memory: the left and right arrow keys now flip its on/off settings, such as Auto-memory - Improved
/skillnames typed mid-message: Claude is now told they are skills, includingdisable-model-invocationones - Improved the contrast of the prompt input border in light themes and of the ❯ before your earlier messages
- Improved delivery of files Claude sends from cloud sessions and Remote Control: an upload that fails on a timeout, a network error or a 502, 503 or 504 is now retried once
- Improved what Claude says when a file cannot be sent for a reason that may be temporary: it now mentions that you can ask for the file again in a few minutes
- Improved the prompt for a held message from another session to show the message between dashed lines, matching other permission prompts
- Improved MCP and other tool permission prompts to show the tool call between dashed lines, matching file edit prompts
- Improved MCP startup in headless mode: a remote server whose first connect fails transiently is now retried without waiting for the slowest server to finish connecting
- Improved files Claude sends from a remote session: large files now stream from disk instead of being read into memory, and a file over the size limit is refused with the server's limit named
- Improved the explanation Claude gives when the server refuses a file it sends from a Remote Control or cloud session, such as an oversized image
- Improved handling of large MCP tool results: less memory, smaller session files, and no extra upload to count tokens for results far over the limit
- Windows: Improved Bash tool speed by removing a subshell that ran before every command
- Changed a shell write through a repo-committed symlink onto a sensitive file or out of the working tree to name where it lands and wait for a person, on lines with a
~target too - Changed Opus 4.7+ and Fable to use a 1M context window by default on Bedrock, Vertex, Foundry and the Claude apps gateway, with no
[1m]suffix (CLAUDE_CODE_DISABLE_1M_CONTEXT=1keeps 200K) - Changed replies from
claude agentsto arrive as queued messages; slash commands other than/stopsent while a turn is running now run when it ends - Changed whole-tool
Bashallow rules and allowing hooks to prompt for, not run, shell writes to files Claude Code's file tools refuse outright (the Anthropic profile store, the host credentials file) - Changed right-click paste on Windows and Linux, and middle-click paste on Linux, to happen when the button is released; moving the pointer away before releasing cancels it
- Changed MCP server
alwaysLoad: falseto defer all of that server's tools behind tool search - Changed screen reader mode to write new or changed lines without first pausing with the cursor at the start of the line; set
CLAUDE_AX_PREPARK_MS=50to restore the pause - Changed automatic model switches after a flagged message to keep your current effort level instead of the new model's default
- Changed waiting permission prompts to show oldest first, so a new prompt no longer covers the one you're reading (prompts with a countdown still open on top)
- [VSCode] Added "Run in background" to a running command or sub-agent, to move it to the background and keep working
- [VSCode] Added the output of background shells and Monitors to their cards in the agent map
- [VSCode] Fixed settings dialogs blaming a timeout when Claude Code's reply was too large to confirm a save
- [VSCode] Fixed reopening a cloud session that the side bar already brought to this machine opening it again in a new tab; the side bar is shown instead
- [VSCode] Fixed the side bar's Web tab not listing cloud sessions started after the window loaded; a failed load now says "Remote server is not connected" instead of "No web sessions yet"
- [VSCode] Fixed a tab restored after a reload starting a second Claude process on a conversation the side bar already has open; it now shows the "still open somewhere else" notice
- [VSCode] Fixed tool-row file links, session-list links and two hints showing in plain text
- [VSCode] Fixed a background agent's still-running command showing as failed once the main turn ended
- [VSCode] Fixed a user's own
/usageor/contextcommand opening the extension's dialog instead of running when picked from the command menu - [VSCode] Fixed file links in the plan preview tab doing nothing when clicked; they now open the file like links in chat replies
- [VSCode] Fixed opening a tool's input or output in an editor tab failing with "Timeout waiting after 1000ms" on remote hosts such as WSL when the tab is slow to appear
- [VSCode] Improved the Manage plugins dialog: a failed marketplace add, remove or refresh now says what went wrong
- [VSCode] Changed the Claude in Chrome "Enabled by default" switch to also connect the editor's own sessions, which still ask before browser actions
- [Cloud sessions] Fixed occasional failures to fetch from or push to GitHub when GitHub briefly refused a newly issued access token
- [Claude Tag] Fixed Claude posting a failure warning, such as a spend limit notice, in a Slack thread when a background event like GitHub activity woke it and nobody was waiting on a reply
- [Claude Tag] Fixed Claude Tag's spend limits page in admin settings leaving out recently created and private channels in organizations with many channels
- [Claude Tag] Improved Claude's task list in long Slack threads: background work no longer reposts it as a new message on its own, so people following the thread aren't notified
- [Code Review] Fixed finding comments and their "Why this was flagged" text stopping mid-sentence; they now end on a complete sentence
- [Code Review] Fixed Code Review skipping a pull request after a new push when its review had failed twice on the previous commit; it now reviews the latest commit
- [Code Review] Improved the failed-review card on a pull request whose conversation is locked: it now says the lock blocked the review and that nothing was posted or charged
-
🔗 The Pragmatic Engineer The Pulse: RoR creator sparks new “death of coding by hand” debate rss
Before we start: if you happen to be in San Francisco on Thursday, 5 November, join me on the System Update with The Pragmatic Engineer event. This is an evening with OpenAI, Linear and DoorDash and myself, organized by Sentry. We get into what AI-augmented automations they're running in prod, and how it's going, in an off-the record (that is: not recorded!) and raw conversation. Seats are limited, and you can RSVP here.
Hi, this is Gergely with a bonus, free issue of the Pragmatic Engineer Newsletter. In every issue, I cover Big Tech and startups through the lens of senior engineers and engineering leaders. Today, we cover one out of four topics from the last week 's issue of The Pulse . Full subscribers received the article below seven days ago. If you 've been forwarded this email, you can subscribe here .
The creator of Ruby on Rails, David Heinemeier Hansson, caused quite a stir last week with comments in his Rails World keynote, when he revealed that coding by hand is dead at his company, 37signals.
This is a big deal because 37signals created Ruby on Rails, and they are known for their software craft there, especially when it comes to code quality. It's also a business that's 27 years old and is profitable. Despite that pedigree, DHH caused a stir among the dev community, saying:
"At 37signals, a couple of weeks ago, we made the decision that it clearly means we 're done writing code by hand. We have gone pencils down on the idea that we were gonna write code by hand, as a normal course of business creating things.
Writing code by hand at 37signals is now an exceptional state. It is like seeing a bug in Sentry: something here went wrong; why was the agent not able to produce what we wanted? Okay, maybe for a little while, we'll still get the old pencil out and dot it down for them, but then we fix the machine, we fix the factory, we get things going again. This is a recognition of what's already happening."
DHH compared the maturation of AI tools into being highly capable at coding with the impact upon the craft of painting of the arrival of the camera:
"On November 24th, 2025, we got the "Kodak Brownie" of our era. We got Opus 4.5. AI technology, accessible in a harness that many people could afford to use and experience for the first time what it's like to create software in pairing with a new form of intelligence. This was the tipping point for me. There was everything before November 24th, and then there was everything after. This is going to be the date that history books going forward will mark as the inflection point for the age of agents."
He shared how 37signals has embraced a future where coding by hand is almost entirely absent:
- Embracing native mobile apps instead of web: famously, 37signals is bearish on native iOS and Android apps and has built web versions instead. With AI, they are betting on native apps being much easier to be built with a small team and are already building new ones.
- Moving backend services to Rust, not Ruby : This is due to performance reasons and because agents write good enough Rust. That's remarkable to hear from the creator of Ruby on Rails!
- Ruby on Rails remains for web apps : 37signals is not leaving RoR behind, but only because Ruby on Rails' convention-over-configuration design makes it easy for agents to work with it.
DHH closed by revealing that he no longer even thinks of himself as a professional programmer (emphasis mine):
"I have retired from being a professional programmer. I think it was somewhere around 4 to 5 months ago, maybe March. I spent a quarter of a damn century chiseling code by hand and loving every moment of it. This is not something to look back upon with regret; this is something to look back upon with joy and accept that it is over.
Writing code by hand is no longer an economically productive enterprise for the vast majority of programmers working at the vast majority of companies. On the other side of that is a new career as a professional maker of things, steering intelligence that was only available in science fiction up until a few moments ago.
One of the things we're gonna have to revisit is everything we think we know about software architecture. The main tool that we've used for a very long time is abstractions. Abstractions don't make quite the same sense in the age of agents. The reason we did abstractions was in part not to repeat ourselves; well, now the price of repetition has gone to near zero."
It's worth noting DHH's keynote chose a spicy topic for a conference attended by engineers who are personally and professionally invested in the craft of building software!
Decline of coding by hand is long predicted
In the first issue in The Pragmatic Engineer this year, on 6 January, I wrote:
"When AI writes almost all code, what happens to software engineering? No longer a hypothetical question, this is a mega-trend set to hit the tech industry. (...)
The bad news is that change will probably be rapid. It's barely been a year since the idea of Claude Code was born in Boris Cherny's head, and already similar tools like OpenCode, Codex, Factory, Amp, Cursor, and more capable agents are changing how software is written. Change has always been part of working in tech, but I cannot recall it being this fast, or happening across the whole industry at once!"
I concluded that this change was on its way, based on my own experience of building software with Opus-4.6 and GPT-5.2, and from talking with experienced engineers who had resisted "AI hype" for good reason, but who had come to see that AI can now generate code that's "good enough" in many cases.
Back then, I made a few predictions about what will happen when AI agents are producing most of the code for engineers:
- Sloppier code
- Weak software engineering practices hurting sooner
- "Coders" who are not software engineers see less demand
- Tougher work-life balance for engineers
- Junior engineers pushed to become seniors, fast
- Computer science education increasingly required for new hires
- A massive explosion in code and software, for which someone must be accountable
So far, it's a messy transition and we engineers are responsible and accountable for a lot more code that we didn't write, but which is in production anyway.
Non-engineers also getting into agents
At the end of January, I shared a deepdive that was pretty close to home for me: my brother's 30-person, 15-engineer startup, Craft Docs, made its own sharp pivot to AI by building their own AI harness for non-engineers - called Craft Agents - two weeks before Claude Cowork was released, and months before ChatGPT Work launched.
Craft resisted the temptation to use AI when it did not feel productive, but with the model releases of November 2025, they found LLMs are not only useful for coding, but also for non-engineering work like customer support. In the deepdive, I went into more detail about non-engineering use cases (which engineers enabled) like:
- Automatic triaging of bug reports with agents
- Data enrichments added to all workflows
- Customer support "skills" like processing feature requests
- The marketing team building websites without devs
- HR automating tedious work
- Finance automating personal workflows
Craft Docs seemed early to a trend that has become more widespread, by having both their own engineering and non-engineering folks onboard to an AI harness. Now, there are signs other companies are doing the same: at OpenAI, non- engineering units like finance, recruitment, and legal moved over to Codex in June 2026:
Non-
engineering teams moved over to use OpenAI 's AI harness, Codex (now renamed
to ChatGPT Work.) Source: Inside OpenAI 's software
factoryIn some ways, it could be comforting to know that it's not only software engineering where the tools and workflows are quickly changing: every other function in tech is experiencing the same!
It 's messy right now
Just last weekend, a rant by an anonymous engineer in Big Tech hit a nerve with many people in the industry. An engineer with the username voxium posted (emphasis mine):
"The state of engineering right now is horrible. It has been half a month since I started a new role at a big company.
Nobody knows anything here. The specs, code, tests, PRDs, tickets, resolution of those tickets, reports, etc., everything is made by Claude Code. Nobody on my team likes this.
They are being forced to ship as much as they can. I have heard multiple times from higher management that pushing code is not a bottleneck, so why are we slow?
People are working 12 to 13 hours a day just to press enter. Nobody is reading anything. Humans in corporate are doing nothing on their own.
Everyone, literally everyone, from an L1 to an L7 engineer here is doing the same thing. Talk to Claude.
There is no sense of victory. Nobody is resolving bugs. In reality, nobody is thinking anymore. Everything is done by LLMs. It is so soul-sucking.
I would not mind it, to be honest, if we were at least given the time to check out the code and see what is going where. But no, the goal is to just ship. No matter what happens."
This post rings true because it is happening at many places where there's more AI usage, engineers do "outsource" thinking to LLMs, and end up not caring about anything else except shipping something to production.
Quality in decline
Since the beginning of the year, the quality of software has been degrading pretty much everywhere, much of it caused by over-reliance on AI, or perhaps more accurately, the outsourcing of thinking and decision making to AI. In July, I moved my video podcast off of Spotify after a series of unexplainable outages, and Spotify's engineering team seemed to take no real pride or accountability in fixing the root causes of the issue.
Only this week, Uber shipped a new feature to production in the Uber Eats app - a new way to select extras with your food order - with seemingly no QA testing:
Uber
Eats this week, when I attempting to order a burger. Can you spot two obvious,
sloppy bugs on this page?Inside this new "add-ons selector" in Uber Eats, I noticed three bugs at once:
- " Choose up to 999": no engineer, designer, or PM bothered to check what happens when a restaurant does not fill out a number on how many toppings to add, or adds a ridiculously large number. The most toppings that my screen allowed to be selected was six and not 999, so I could not even choose the option of 999 buns for my burger.
- Sloppy overflow. A rule of thumb, during my time at Uber, was that text will never overflow, even when localized. Basilcummaynaise (basil mayo in Dutch) broke this rule but still shipped.
- Functional bugs in the selector. I originally tried to order from my favorite Mexican place: a bowl with no rice or bulgur as the base. There's the option to select "rice", "bulgur" or "nothing" as the base, but selecting "nothing" counts as an extra side, and the app doesn't allow the ordering of a bowl with no base.
I've used the Uber Eats app for years, and this was the first time I saw such a sloppy feature release. I assume that devs and PMs building it have all "checked out", stopped doing proper QA, and assume that the agent will take care of all of it. There's no other way to explain three bugs shipped to all customers but seemingly noticed by nobody until I posted about it. To the Uber Eats team 's credit, they reached out and are looking into fixing all three issues.
Software engineering to be more important than ever
I'm personally past the shock and grief stages of agents taking over the activity of coding. At first, I assumed this change would reduce the amount of work for engineers. But, counter-intuitively, that actually seems to be growing:
- We need to understand the characteristics of LLMs better. LLMs feel familiar as they can produce text in a way only humans could do before. But they are less reliable, still prone to hallucination, suffer from capability gaslighting, and many other problems. They can also be expensive and slow.
- New systems need to be engineered. Agentic "software factories" can now produce code, based on the input provided. But how is this code validated? How much can detecting defects or various issues be automated? This is a brand new area, and we need to build new types of systems, often based on old ideas. One such example is OpenAI's software factory, the other one is Ramp's Inspect internal coding harness.
- Nondeterministic LLMs can generate deterministic code. An area I feel is under-explored and under-appreciated is the use of LLMs to substitute LLMs usage in agentic "software factories" with deterministic code they generate. For example: instead of running AI code review that is expensive and slow on all PR requests, could AI generate linters that catch the majority of common issues? If this is possible, complex lint rules would run faster, be more reliable and cheaper to execute than LLM calls.
New categories of systems and products will be built by engineers who "get" LLMs and AI engineering. We are seeing the majority of venture funding pour into AI companies because AI creates new business models, new revenue streams, and disrupts "traditional" software. For example, who would have thought that companies would spend tens of thousands of dollars, per engineer, on AI coding tools? Or that the category of AI inference providers would become as massive as it already is from barely existing two years ago?
This technological change will re-jig parts of the tech industry: the winners will surely win big, and teams and companies choosing inaction could be out- executed and displaced by nimble competitors. And in many ways, this is great news for us software engineers who keep up with the technology. Companies are now investing in innovation and are willing to pay top-of-market for software engineers who can help them build AI products or become AI-native.
Read the full issue of The Pulse this is from, or check out this week 's The Pulse. This week's issue covers:
- Firebase: global outage & poor handling by Google. The Firebase iOS SDK crashed after a backend change, crashing all apps which used Firebase analytics for 2-6 hours. Google did not update the status page or offer any postmortem, which is a head-scratcher from a company known for standout incident management practices.
- OpenAI 's platform play from AWS playbook? OpenAI is becoming a platform where it's possible to allocate ChatGPT spend on open models and AI offerings from among 16 partners, not just OpenAI models. It's fair to ask if Anthropic will consider a similar platform play.
- More data on companies moving to open models. Vercel's AI gateway shows 60% of model spend goes to open weight models, and OpenRouter also shows open models are being more used than closed ones.
- Why CTOs and VPEs are quitting en masse: another take. What if it's not "founder mode", but about people who love building software, feeling like they can do it solo (or with a small team) with AI tools?
-
🔗 @HexRaysSA@infosec.exchange The IDA 9.5 Beta is live! mastodon
The IDA 9.5 Beta is live!
◾ IDA MCP + Assist for agentic RE
◾ New TriCore, Hexagon and DEX/ODEX decompilers
◾ Recursive decompilation
◾ New Malware Analysis add-on
◾ And more...Beta members: it's in your Download Center now.
Not in the Beta program? Join from the customer portal today.
-
🔗 Hex-Rays Blog IDA 9.5 Beta is available rss
IDA 9.5 Beta is now available
The IDA 9.5 Beta is live, starting today. If you're part of our Beta Program, the new build is already waiting in the Download Centerof your customer portal, so you can put it to work right now. Not enrolled yet? You can join the program in a few clicks from your customer portal dashboard, and your testing is what tells us a feature is ready for production, where a regression slipped in, and what still needs sharpening before release.

-
🔗 r/Harrogate Moving to Harrogate need advice rss
Hey everyone moving from Manchester what area is nice I don't drive so need to be around the shops and a ideal world near a bus route and a mosque any idea what area I should I look at
submitted by /u/AlifanofmalcomX
[link] [comments] -
🔗 r/Harrogate Harrogate in it's glory days rss
| submitted by /u/LowGuide2746
[link] [comments]
---|--- -
🔗 exe.dev ETOOMANYTHINGS? Run Fewer Agents rss
At a conference last week, I sat through a bunch of software factory demos, and they were all about task management. Kanban boards, Slack interfaces, email interfaces, dependency graphs, ticketing, graphical interfaces, you name it.
Why is everyone building task management and calling it a software factory? Agents are slow to execute. And the obvious, easy fix to latency is to hide it by starting a new agent every time you get blocked.
But concurrency is really rough on humans. It’s stressful. It trashes flow state and thrashes our mental page caches.
Task management isn't a solution; it's a band-aid.
This is a call to arms. Let’s make it possible to be equally productive with fewer agents.
Bitter Lesson to the Rescue?
As models improve, “good enough” models will get ever faster. This will naturally reduce latency, and in turn concurrency. We’ll still have to do some old-fashioned engineering, like making tests run fast.
But we don't have to wait! Here are some things we’ve experimented with at exe.
Use a Fast Model for Talking With Humans
Start tasks with a powerful model. But instead of reading a wall of text and writing an essay in response, put all the comms in the hands of a fast, competent model like Luna 6. Give Luna no coding tools, and make its context window intentionally short and focused on what the human conversation requires. Then, when you go to check in on a task, you can have real-time discussions with Luna. With a fast feedback cycle, attention doesn’t wander, and you can stay engaged for longer. Eventually, Luna exhausts its ready information, and the beefier model takes back over, with rich, substantial user feedback to work from.
This works reasonably well. But we found that, shock of shocks, absorbing information via chat was rather constraining. It was frustrating to have communications be squeezed through a tiny pipe. Luna might be an exciting, bendy straw, but it’s still a straw.
Screen Recording
That led us to focus on the HCI aspects of the problem. Back when I coded by hand, I would stare intently at screens full of text, processing it slowly, navigating freely between files, thinking. But the UI that is presented by most coding harnesses is a single linearized text thread, typically interspersed with lots of irrelevant tool call noise. There’s very little human agency or control, and no organization. You can’t even rely on the most important content being at the bottom: noise from straggler subagents often drowns out the agent's primary response.
So: How can we restore human agency and optimize for human I/O, rather than making the human adapt to the agent?
For input, the human retina is a powerful information processing system, when given structure to work from. That is, not a wall of text. I can still pick a Go panic stacktrace out of terminal logs scrolling by at 30fps.
For output, even for the fastest typists, typing is typically slower than speaking. Also, typing requires coordination with lots of other parts of the computer. The text input window has to be focused. Your cursor must be in the right place. It constrains the other things you can do with your keyboard and mouse and what you can look at. Many coding agents, including Shelley, worked around this with affordances for annotating text, diffs, and web pages, but it still requires clicking, and it fundamentally requires cooperation from everything else.
We recently launched a simple show-and-tell feature in Shelley that is tailored to humans. Click the camera button. Shelley records audio so that you can think out loud as you explore, and it records your screen. Anything on screen is a thing that you can point at, talk about, refer to. Poke around, muse, backtrack, trail off, resume, change your mind, whatever. Agents are extraordinarily good at understanding rambling.
When you stop recording, Shelley transcribes all of your audio, including word-level timestamps, and makes a contact sheet from the video. It correlates your words with what was on screen when you said it. The computer does the hard work of collating all of the feedback.
In my experience, this form of interaction is typically deeper, richer, longer, and more engaged than chatting with an agent. This is particularly useful for UI work or anything visual.
DOM Recording
What if you're not working on a visual task? The same ideas can be adapted nicely to other forms of engineering work.
Another feature of human cognition is that we work well with concrete examples, not abstract descriptions. Agents are happy to be told, but humans prefer to be shown. Humans also benefit from diagrams, charts, and visually structured diffs.
With this in mind, I have been using a slightly different system that we haven't shipped in Shelley (yet). Instead of spewing all its output to me in raw text, the agent generates an HTML artifact. This HTML is standardized to reduce visual noise, ruthlessly reduce verbosity and complexity, structure information visually, and make navigation easy. Typical sections include notable decisions made, open questions, examples, and git diffs.

This HTML artifact includes a record button. When clicked, it records audio. Instead of using screen recordings, it uses a cheaper, more precise in-page DOM tracker JavaScript that tracks what is being viewed, scroll position, where my pointer is (if there is one), what text is highlighted, where I click/tap, all with timestamps. As before, it correlates timestamps with audio and feeds it all back to the agent.
I find that this reduces mental overhead; I am freed to focus on the content.
More Latency Hiding
Because really engaging with content takes time, there is an additional latency-hiding trick one can pull. I am still experimenting with this, but the idea is to take prefixes of my feedback while it is still being recorded, send them to the agent, and livestream its responses back to a tab on the very same HTML artifact I'm already looking at. By the time I’ve spent 10 minutes reading and probing, I will inevitably have asked questions and provided provisional guidance. Before I context switch away, some of my questions have answers waiting. Some of my design decisions have follow-up questions. And the camera is still rolling. This sneaks in an extra round trip with the agent without breaking my concentration.
There’s Still Work to Do
I am not yet at a point where I have the same net throughput with one or two agents that I do with many, but my attention is being restored to me, little by little.
-
🔗 pydantic/monty v1.0.1 release
release: 1.0.1 (#957)
-
🔗 smol-machines/smolvm smolvm v1.22.1 release
What's Changed
- Bump libkrunfw to the guest kernel with conntrack marks so nft ct mark rules load by @BinSquare in #1493
- Stop the workload wait ticker without waiting out its sleep by @LoganGrasby in #1495
- Pin libkrunfw to the merged conntrack-mark commit on main by @BinSquare in #1494
- Restrict serve mTLS clients by certificate subject CN by @BinSquare in #1499
- Speed up restored starts with --keep-identity and RAM prefetch by @BinSquare in #1497
- Preserve HTTP machine records when storage cleanup fails by @hfiguera in #1498
- Give machine update the --outbound-localhost-only flag create has by @BABTUNA in #1500
- Prefetch restored RAM and allow keep-identity restores in the embedded runtime by @BinSquare in #1502
- Bump the workspace to 1.22.1 by @BinSquare in #1496
New Contributors
Full Changelog :
v1.22.0...v1.22.1 -
🔗 backnotprop/plannotator v0.27.24 release
Follow @plannotator on X for updates
Missed recent releases? Release | Highlights
---|---
v0.27.23 | Bitbucket Cloud PR review, Question UI for answering agents in place, opt-in auto-update, viewed files remembered, review another repo or worktree
v0.27.22 | Plans open the docs they link to, Claude review jobs locked down, Code Tour on Linux without Claude's sandbox, Pi reviews the latest plan
v0.27.21 | Remote and phone sessions load several times faster, real Request changes on GitHub, model pickers show real names, OpenCode fixes
v0.27.20 | Mistral Vibe support, annotate gets the full Options menu and Settings, jj Commits panel, long lines wrap in plan code blocks
v0.27.19 | Before/After image previews in code review, file comments as GitHub file threads, forge-correct#123links,/plannotator-lastfinds the right session
v0.27.18 | Model pickers from your installed Claude and Codex (Opus 5.5, Fable 5.1, GPT-6), unsent PR review comments survive new pushes
v0.27.17 | Diagram files open in the diagram viewer, OpenCode switches model with agent, idle review stops polling the git remote, Tree is the default review view
v0.27.16 | Themed diagrams on Mermaid 12, comment on any node or edge, patch-file review, embedded HTML documents render
v0.27.15 | Plannotator TUI and Herdr Annotate announcement, element context on pinpoints, HTML links open as linked documents, All files panel, Classic diff default
v0.27.14 | Pi plan progress survives compaction, Codex threads across rollout files, WSL browser setting, Mod+E edit mode
v0.27.13 | Open a review on a specific base (--base,--diff-type), symlink containment on /api/doc, CI flake fix, Amp decision relayWhat's New in v0.27.24
A small fix release for two code review problems found while recording a walkthrough of v0.27.23.
Image previews stay in the all-files view
When a diff changed images, the Before/After previews in the all-files view disappeared after a short scroll, and the view jumped ahead to the code files. The list measured each image card as having no height, so it dropped the cards as soon as they left the top of the screen. The cards are now measured at their real size, so they stay in place and keep their height while you scroll up and down. Scrolling a large diff is as fast as before.
(#1652)
PR comment previews open on the commented line
In the PR Comments panel, the small code preview under a comment showed the start of the changed block, not the line that was commented on. For a new file that meant lines 1 to 5, even when the comment was on line 28. This was most visible on Bitbucket, but long blocks on GitHub had the same problem. The preview now opens on the commented line and the three lines above it, or on exactly the lines of a range comment. "Show full context" still shows the whole block.
(#1651)
Install / Update
macOS / Linux:
curl -fsSL https://plannotator.ai/install.sh | bashWindows:
irm https://plannotator.ai/install.ps1 | iexClaude Code Plugin: Run
/pluginin Claude Code, find plannotator , and click "Update now".Pi: Update
@plannotator/pi-extensionto 0.27.24 and restart Pi.OpenCode: Clear cache and restart:
rm -rf ~/.bun/install/cache/@plannotatorWhat's Changed
- fix(review): PR comment previews open on the commented lines by @backnotprop in #1651
- fix(review): keep all-files image previews in the virtual list by @backnotprop in #1652
Full Changelog :
v0.27.23...v0.27.24 -
🔗 HexRaysSA/plugin-repository commits sync repo: +2 releases, -1 release rss
sync repo: +2 releases, -1 release ## New releases - [ida-mcp](https://github.com/hexrayssa/ida-mcp): 20260930.0.1 - [ida-nexus](https://github.com/hexrayssa/ida-nexus): 0.13.1 ## Changes - [ida-codemode](https://github.com/hexrayssa/ida-codemode): - removed version(s): 0.6.0 -
🔗 Rust Blog Announcing Rust 1.99.0 rss
The Rust team is happy to announce a new version of Rust, 1.99.0. Rust is a programming language empowering everyone to build reliable and efficient software.
If you have a previous version of Rust installed via
rustup, you can get 1.99.0 with:$ rustup update stableIf you don't have it already, you can get
rustupfrom the appropriate page on our website, and check out the detailed release notes for 1.99.0.If you'd like to help us out by testing future releases, you might consider updating locally to use the beta channel (
rustup default beta) or the nightly channel (rustup default nightly). Please report any bugs you might come across!What's in 1.99.0 stable
extern "C" variadics
Rust 1.99.0 stabilizes defining C-ABI variadic functions with "C" and "C-unwind" ABIs. Variadic functions defined this way use a variable argument list (
...) and accept an arbitrary number of arguments. Rust could already call externally-defined variadic functions (e.g.,libc::printf). With Rust 1.99, these functions can now be written in Rust itself:/// SAFETY: must be called with (at least) 2 i32 arguments. unsafe extern "C" fn sum(mut args: ...) -> i32 { // SAFETY: guaranteed by the caller. let a = unsafe { args.next_arg::<i32>() }; let b = unsafe { args.next_arg::<i32>() }; a + b } fn foo() -> i32 { unsafe { sum(0i32, 2i32) } }The type of
...isVaList, which is ABI-compatible with the Cva_listtype across targets. What types can be read from aVaListis guarded by theVaArgSafetrait.For more details on c-variadic functions, see the Reference. This release also stabilizes support for defining naked variadic functions with non-"C" ABIs, which must be written via inline assembly.
Layout information from raw pointers
This release settles the safety requirements for retrieving the size and alignment on raw pointers to both
Sized(trivially safe, already possible on stable) and non-Sizedtypes.This is done by stabilizing three functions:
Recommend against round-trip unleaking after
Box::leakWhile there are no changes to the language semantics in Rust 1.99, we have updated the documentation on
Box::leakto recommend against patterns that later deallocate that memory. This was done because such code was found to have problematic interactions with current and future potential compiler optimizations, and is especially problematic with the upcoming stabilization of custom allocators. Instead,Box::into_non_nullorBox::into_rawshould be preferred.This guidance also applies to other
leakfunctions in the standard library.Stabilized APIs
IntoIteratorforBox<[T; N]>IntoIteratorfor&Box<[T; N]>IntoIteratorfor&mut Box<[T; N]>VecDeque::retain_backcore::ffi::VaListBox::into_non_nullBox::from_non_nullVec::into_partsVec::from_partscore::mem::size_of_val_rawcore::mem::align_of_val_rawcore::alloc::Layout::for_value_rawString::from_utf8_lossy_ownedstring::FromUtf8Error::into_utf8_lossyFusedIterator for StepBy<I>std::fs::set_timesstd::fs::set_times_nofollow
Other changes
Check out everything that changed in Rust, Cargo, and Clippy.
Contributors to 1.99.0
Many people came together to create Rust 1.99.0. We couldn't have done it without all of you. Thanks!
-
🔗 Console.dev newsletter tinyjs rss
Description: Native webview desktop apps.
What we like: Not Electron, so no bundled browser - uses the system native webview (macOS WebKit, Windows WebView2, Linux WebKitGTK). Native chrome with system API access e.g. files, sockets, processes. Small app bundles (runtime is 6MB). Everything is HTML inside.
What we dislike: Everything is HTML, so whilst the wrapper feels native, the app isn’t.
-
🔗 Console.dev newsletter celld rss
Description: Self-hosted durable objects.
What we like: Each object (server) gets its own isolated SQLite database. Supports function execution, k/v, queues, cron triggers. Backed by object storage. Compatible with Cloudflare primitives and config. Can be distributed, with state managed in storage.
What we dislike: Still early, with limited fleet management functionality e.g. auto-scaling, etc.
-
🔗 Filip Filmar rules_openxc7: Open-Source 7-Series FPGA Toolchains in Bazel rss
rules_openxc7brings open-source FPGA synthesis and place-and-route for AMD 7-series chips into Bazel. It drives Yosys, nextpnr-xilinx, Project X-Ray, and openFPGALoader through the same target interface asrules_vivado. This post covers the toolchain architecture, how the rules translate Xilinx constraints, reproducible bitstreams, and how to swap between Vivado and open tooling in one line.Why open tooling in Bazel
Building for FPGAs usually means one of two extremes. You either install a 100-gigabyte proprietary vendor suite, or you maintain fragile shell scripts around open-source tools.
rules_openxc7avoids both problems. -
🔗 New Music Releases Nightwish - Tribal (live in Amsterdam 2022) rss
Nightwish - a new release is available:
- 2026-10-01: Tribal (live in Amsterdam 2022) (Single)
Amazon: Canada | Deutschland | France | United Kingdom | United States
Visit muspy for more information.
-
- September 30, 2026
-
🔗 Quarkslab's blog Cato VPN Client: Split-Tunnel and Privilege Escalation (CVE-2026-10739) rss
Introduction
During a Purple Team engagement, we had to list and rank risky components across the network. One of them caught our attention: Cato Client, a VPN client program. It was installed everywhere. It runs privileged services. It talks to a GUI. It handles network configuration. From an attacker perspective, this is exactly the kind of software you want to understand. However, saying "this looks risky" is not enough. There is nothing better than
PoC||GTFO. So the question was simple: can we find a real bug and turn it into something useful? This is how it started. And then my teammate sent me this message:💬 YV: "I have a new target for you. It has everything you like: named pipes, protobuf and .NET."
He knows me well. That sounds like a fun challenge.
This writeup covers a local privilege escalation in Cato Client. The bug starts in the split-tunnel upload feature, goes through a named pipe, and ends up with a
SYSTEMshell.Cato Networks SDP Client for Windows is vulnerable to Local Privilege Escalation.
Confirmed on versions
6.2.0and6.4.6on Windows. From vendor:any version before 6.12.6Cato & Split Tunnel
Cato Client is the endpoint client used to connect users to Cato Cloud.
One feature is split tunneling. Cato documentation explains that administrators can let users upload a text file to decide which IP ranges are included or excluded from the encrypted tunnel. The file contains an
includeorexcludemode, followed by IP ranges and masks.In the client UI, this feature appears as a local upload for the split tunnel configuration.

Cato Client settings - split tunnel file upload.A minimal valid
.ccstfile looks like this:include 10.10.10.0/24You can find more information about the split tunnel here.
Walkthrough
Discovery
The first useful observation came from a normal upload.
When a valid CCST file is uploaded, the client accepts it and the (privileged) backend creates split-tunnel state (
.stp) underProgramData.
Valid CCST upload.So I tried the opposite: upload a file with invalid CCST content.
The parser rejects it. That part is normal. But the cleanup path is more interesting: the generated split-tunnel file is deleted.

Invalid CCST upload. The generated file is cleaned up.At this point, it is not yet a vulnerability. A privileged process deleting its own temporary file is not enough. Moreover, the destination directory has sane permissions. A limited user cannot just drop or replace things there:
PS C:\ProgramData\CatoNetworks\SDPClient> icacls .\ST\ .\ST\ AUTORITE NT\Systeme:(I)(F) AUTORITE NT\Systeme:(I)(OI)(CI)(IO)(M,WDAC,WO,GR,GW,DC) BUILTIN\Administrateurs:(I)(F) BUILTIN\Administrateurs:(I)(OI)(CI)(IO)(M,WDAC,WO,GR,GW,DC) BUILTIN\Utilisateurs:(I)(R) BUILTIN\Utilisateurs:(I)(OI)(CI)(IO)(R,GR)As there was no easy win in the filesystem permissions, the primitive was not "just write a symlink in ProgramData and win".
The next question was: how does the low-privileged GUI ask the
SYSTEMprocess to parse and delete those files?Named pipe time
Our internal named pipe tool showed the answer: local IPC over a named pipe.

Cato GUI talks to the service through a named pipe.The pipe is:
\\.\pipe\cato-VPNThe service-side process observed during the proof was:
C:\Program Files (x86)\Cato Networks\Cato Client\winvpnclient.cli.exeand it runs as:
NT AUTHORITY\SYSTEMNice. A low-privileged process asks a
SYSTEMprocess to parse a local file and delete cleanup artifacts.But direct access failed. We couldn't interact directly with the pipe. Sending commands from a random process to the pipe did not work.
With the help of Ghidra, we can see the named-pipe server checks the process talking to it:
GetNamedPipeClientProcessId(...); GetNamedPipeClientSessionId(...);The interesting part is what happens in the client-connected callback after that.
The callback first checks if IPC certificate validation is enabled. If yes, it resolves the executable path from the connecting PID, then verifies the executable certificate:
cVar2 = FUN_1401e5f20(...); // certificate check enabled? if (cVar2 != '\0') { FUN_140767200(local_88, ..., client_pid); cVar2 = thunk_FUN_140753b80(local_88, expected_cert, ...); if (cVar2 == '\0') { "%s: Failed to register client. Could not verify client certificate. Close connection."; FUN_14076fc50(...); // close pipe } }FUN_140767200is the part resolving the client image path:hProcess = OpenProcess(0x1000, 0, client_pid); QueryFullProcessImageNameW(hProcess, 0, local_248, &local_294);The verification helper then parses the executable signature and compares certificate material:
FUN_1407589d0(...); FUN_140758aa0(...); ... iVar2 = memcmp(_Buf1, expected_cert_material, size);Conclusion: a random process does not pass this check. The pipe connection needs to originate from a process image signed with the expected Cato certificate, unless certificate checking is disabled by configuration.
Bypassing the certificate check
This part seemed familiar. In the K7 writeup, the patch also tried to block random clients from talking to a privileged named pipe. Manual mapping a payload into a signed/trusted process was enough to get back inside the IPC path.
Same idea here.
The PoC starts a signed Cato process:
C:\Program Files (x86)\Cato Networks\Cato Client\CatoClient.exeThen it manually maps a DLL payload into it. The DLL connects to:
\\.\pipe\cato-VPNFrom the service point of view, the connection comes from a Cato-signed executable, so the certificate check is not a problem anymore.
Protobuf, you said?
The boring part was rebuilding the Protobuf messages required to talk to the service correctly. This is where AI was helpful: not to find the bug, but to automate repetitive message-building work once the protocol fields were known.
The useful part came from the decompiled .NET client/common code. The Cato GUI ships generated Google.Protobuf classes under
CatoCommon, and the service communication wrapper shows how the real client builds messages.From
ServiceCommunication.cs:private void sendUiRegisterCommand() { ClientToService clientToService = CreateBasicMessage(ClientToService.Types.Commands.UiRegister); clientToService.UiRegister = new C2sUiRegister { UiProcessId = _processId, UiSessionId = _sessionId, UserSidString = _userSidString, AadUserUpn = _aadUserUpn }; SendCommand(clientToService); } public void UploadSplitTunnelFile(string filePath, bool enabled) { ClientToService clientToService = CreateBasicMessage(ClientToService.Types.Commands.SplitTunnelUpload); clientToService.UploadStFile = new C2sUploadSplitTunnelFile { Filepath = filePath, Enabled = enabled }; SendCommand(clientToService); }The generated protobuf classes then give the exact field layout.
From
C2sUiRegister.cs:public const int UiProcessIdFieldNumber = 1; public const int UiSessionIdFieldNumber = 2; public const int UserSidStringFieldNumber = 3;From
C2sUploadSplitTunnelFile.cs:public const int EnabledFieldNumber = 1; public const int FilepathFieldNumber = 2;And
ClientToServicegives the oneof body fields and command IDs used by the PoC.From
ClientToService.cs:public enum BodyOneofCase { UiRegister = 25, UploadStFile = 37, EnableLocalSt = 38, } public enum Commands { UiRegister = 34, SplitTunnelUpload = 49, LocalSplitTunnelEnable = 50, }So the actual work was mostly: extract the protobuf structure from the .NET code, identify the command IDs and body fields, then rebuild only those messages in the injected payload. The C payload source code is available here: CatoPipePayload.c.
The bug
After solving the pipe connection issue, the real bug is simple.
As seen before, the
.stpfile is built using a user SID. However, this value is not derived from the real client token. It can be manipulated by the user inside the named pipe message. The problem: this client-supplied SID is used as part of a filesystem path.The generated split-tunnel output path looks like this:
<split-tunnel-storage>\ccst_<client supplied SID>.stpThat means the SID is treated as two things at the same time:
- identity material;
- a safe filename component.
It is attacker-controlled, and
..\\path components are not rejected.So instead of a real SID, the payload can register with something like this:
\\..\\..\\..\\..\\..\\..\\tmp\\foobarThe service then builds a path that still starts in the split-tunnel directory, but Windows normalizes the traversal and the final write reaches:
C:\tmp\foobar.stpAt this stage we have two related primitives:
- with a valid CCST file: arbitrary file write as SYSTEM, with a forced
.stpsuffix ; - with an invalid CCST file: delete of the generated
.stppath asSYSTEM.
Obviously, the delete is the fun one.
From delete to SYSTEM
The cleanup path is reached when the uploaded CCST file is readable but invalid. To summarize, as a low-privileged user, the flow is:
- signed CatoClient.exe process origin
- cato-VPN named pipe certificate gate passes
- UiRegister with attacker-controlled UserSidString
- UserSidString reused in ccst_.stp
- path traversal escapes the split-tunnel directory
- SplitTunnelUpload points to an invalid CCST file
- parser fails
- cleanup deletes the generated .stp path as SYSTEM
The last "problem" (which is not really a problem) is that we control the path for the delete operation, but the generated target has to end with
.stp. Solution? Symlink, of course. We only need to redirect this privileged delete to something useful.The PoC uses the well-known
C:\Config.MsiWindows Installer rollback technique. The idea is documented by ZDI and was also used in previous file/folder delete LPE chains.The Cato-specific redirection is:
C:\poc -> \RPC Control \RPC Control\deleteme.stp -> \??\C:\Config.MsiThen the malicious SID points the generated
.stppath to:C:\poc\deleteme.stpWhen Cato cleanup deletes that path as
SYSTEM, it reaches:C:\Config.MsiProcmon evidence showed:

After that, the Windows Installer rollback chain does the rest and gives code execution as
SYSTEM.The autonomous proof of concept
The final PoC is an autonomous wrapper:
It embeds:
- the MSI rollback payload;
- the rollback files;
- the Cato pipe payload DLL.
Expected flow:
- Prepare the Windows Installer rollback state under
C:\Config.Msi. - Create the
C:\poc\deleteme.stp -> C:\Config.Msiredirection. - Manual-map the embedded Cato pipe DLL into signed
CatoClient.exe. - Trigger split-tunnel upload cleanup with invalid CCST input.
- Wait for
C:\Config.Msito be deleted bywinvpnclient.cli.exeasSYSTEM. - Run the second MSI stage and launch the configured command.
And the result:

First LPE PoC.
SYSTEM shell after the Cato delete and MSI rollback chain.Full video with a self-contained binary: CatoEOP.mp4
Conclusion
This one was fun because it was not a single obvious bug. The service did have a client check. The directory permissions were not broken. The cleanup delete only happens on invalid
.ccstfile upload. The named pipe was not open to everyone. But the chain was there:- signed process origin
- SID path traversal over protobuf
- invalid CCST cleanup
- SYSTEM delete operation
- Symlink redirection
- SYSTEM shell
This is exactly why I like turning risk assessment into exploitation. "This software is installed everywhere and runs privileged code" is a useful warning. But a working 0day LPE demonstration is much better. It proves the risk is not only a line in a report. I also have to say a word about AI in this one. It wasn't useful to find the bug, but it saved me a lot of time, probably a few days, to build a working autonomous PoC based on previous work.
Disclosure timeline
Below we include a timeline of all the relevant events during the coordinated vulnerability disclosure process with the intent of providing transparency to the whole process and our actions.
- 2026-06-02 : Quarkslab reported the vulnerability to Cato Networks.
- 2026-06-03 : The vendor acknowledged our report.
- 2026-06-15 : Quarkslab requested an update, the vendor replied on the same day that the vulnerability was pending remediation by the engineering team.
- 2026-08-06 : Quarkslab requested a tentative timeline for remediation.
- 2026-08-06 : The vendor informed that the vulnerability was already fixed internally and the rollout was expected at the end of August. Asked Quarkslab to hold off disclosure until then.
- 2026-09-16 : Quarkslab asked for a status update and a clear release date too coordinate disclosure.
- 2026-09-29 : Quarkslab asked the vendor if a CVE was assigned to the vulnerability.
- 2026-09-30 : The vendor informed that CVE-2026-10739 was assigned to the vulnerability.
- 2026-09-30 : The vendor published a Security Announcement to customers.
- 2026-10-01 : This blog post is published.
References
- Cve.org - CVE-2026-10739
- Cato Networks - CVE-2026-10726 & CVE-2026-10739 that Impacts Windows Client Versions Lower than 6.12.6
- Cato Networks - Routing with the Cato Client / Split Tunnel Policy
- GitHub - CVE-2025-48799 (used as a main reference for PoC automation, thanks!)
- PoC payload source - CatoPipePayload.c
- Quarkslab - K7 Antivirus: Named pipe abuse, registry manipulation and privilege escalation
- Quarkslab - Avira: Deserialize, Delete and Escalate - The Proper Way to Use an AV
- ZDI PoC - FilesystemEoPs / FolderOrFileDeleteToSystem
- ZDI - Abusing Arbitrary File Deletes to Escalate Privilege
-
🔗 backnotprop/plannotator v0.27.23 release
Follow @plannotator on X for updates
Missed recent releases? Release | Highlights
---|---
v0.27.22 | Plans open the docs they link to, Claude review jobs locked down, Code Tour on Linux without Claude's sandbox, Pi reviews the latest plan
v0.27.21 | Remote and phone sessions load several times faster, real Request changes on GitHub, model pickers show real names, OpenCode fixes
v0.27.20 | Mistral Vibe support, annotate gets the full Options menu and Settings, jj Commits panel, long lines wrap in plan code blocks
v0.27.19 | Before/After image previews in code review, file comments as GitHub file threads, forge-correct#123links,/plannotator-lastfinds the right session
v0.27.18 | Model pickers from your installed Claude and Codex (Opus 5.5, Fable 5.1, GPT-6), unsent PR review comments survive new pushes
v0.27.17 | Diagram files open in the diagram viewer, OpenCode switches model with agent, idle review stops polling the git remote, Tree is the default review view
v0.27.16 | Themed diagrams on Mermaid 12, comment on any node or edge, patch-file review, embedded HTML documents render
v0.27.15 | Plannotator TUI and Herdr Annotate announcement, element context on pinpoints, HTML links open as linked documents, All files panel, Classic diff default
v0.27.14 | Pi plan progress survives compaction, Codex threads across rollout files, WSL browser setting, Mod+E edit mode
v0.27.13 | Open a review on a specific base (--base,--diff-type), symlink containment on /api/doc, CI flake fix, Amp decision relay
v0.27.12 | Unified decision control, token hover cards, local-vs-remote diff, approval notesWhat's New in v0.27.23
This release brings Bitbucket Cloud to PR review, lets agents put questions in a plan that you answer in place, and adds an opt-in setting that keeps Plannotator up to date for you. Code review also remembers which files you have viewed, can be pointed at another repository or worktree, and several fixes land for Pi and keyboard users. It contains thirteen pull requests, three of them from community contributor @oorestisime.
Bitbucket Cloud pull requests
PR review used to recognize only GitHub and GitLab links.
plannotator review https://bitbucket.org/<workspace>/<repo>/pull-requests/<id>now opens a Bitbucket Cloud pull request with the same review experience:- the diff and a local checkout for full file access
- existing comments, with their code context
- image previews
- Guided Review and Code Tour
From the review you can post general and inline comments, approve, or request changes. Bitbucket has no command-line tool like
gh, so Plannotator talks to its API directly with an Atlassian API token. SetPLANNOTATOR_BITBUCKET_EMAILandPLANNOTATOR_BITBUCKET_TOKEN; the docs list the scopes the token needs.Under the hood, PR mode now runs through one provider per platform. The screen shows only what that platform supports, which is why the PR Artifacts panel is absent on Bitbucket. GitHub and GitLab were moved onto the same seam and behave exactly as before. We checked this by recording every
ghcall against real stacked PRs on the old and new code and comparing them.Not yet supported on Bitbucket:
- replying to or resolving existing threads
- syncing viewed files back to Bitbucket
- stacked PRs
- Bitbucket Data Center
(#1641, toward #1583 and #671, requested by @Toparvion and @bahayden)
Questions in plans, answered in place
An agent can now put questions in a plan or document as
:::questionblocks with checkbox choices and aRecommended:line. Plannotator draws each one as a card. You can:- pick an answer, type your own under "Other…", add a note, or skip it
- take the agent's suggestion with Accept recommended
The header shows how many you have answered, and clicking that count jumps to the next open question. The sidebar lists them all.
Answers travel with your feedback in an "Answers to your questions" section at the top. If answering questions is all you did, the main button reads Send answers. The agent is then told you answered its questions and asked to revise the plan, rather than being told the plan was rejected. Answers are saved like any other annotation, so they survive a reload and can be undone. Question and option text stays selectable, so you can still comment on the wording.
The
/plannotatorskill teaches agents the syntax, including when not to ask. A document with no questions renders and exports exactly as before. If you already used:::questionas an ordinary callout, it now renders as an answerable card. One limit to know: approving a plan in Claude Code passes no message to the agent, so answers given with Approve do not reach it. Plannotator warns you before that happens.Opt-in automatic updates
The "new version available" notice tends to appear while you are busy with something else. You can now turn on Keep Plannotator up to date in Settings, with
"autoUpdate": truein~/.plannotator/config.json, or withPLANNOTATOR_AUTO_UPDATE=1.With it on, Plannotator checks for a new release at most once a day, in the background, after a session has already started. It never slows down a hook or a command. When a newer version exists and no other review is open, it runs the normal install script for that exact version and writes the output to
update.log. The next time you open Plannotator it tells you what version you are on now, or that the update failed and where the log is.The installer now remembers the flags you installed with, such as
--minimalor--skip-codex, so an automatic update installs the same way you did. The install scripts also replace the binary safely while it is running: on macOS and Linux the new file is renamed into place, and on Windows the running copy is moved aside first. Auto-update is off unless you turn it on, and it skips development builds, the OpenCode and Pi packages, and binaries in folders it cannot write to.(#1636, closing #1634, requested by @drj613)
Code review remembers what you have viewed
Files you mark as viewed now stay viewed the next time you review the same branch, including after a restart or after sending feedback. A mark is tied to the exact contents of the file you saw, so a file that changed since then comes back unviewed. This works for local Git reviews and pull requests. Each mark is a small record in
~/.plannotator/review-progress/that holds the file's path. Turn it off withPLANNOTATOR_REVIEW_PROGRESS=0or{ "reviewProgress": false }, andplannotator uninstall --purgeremoves it.Viewed marks are now kept per diff type. Viewed marks stored in older drafts are ignored once, the first time you open a review after updating.
(#1632, closing #1136, by @oorestisime)
Review another repository or worktree
plannotator review ../other-repoand/plannotator-review ./worktrees/featurenow review that directory without moving your agent session:- A folder inside a repository reviews the whole repository.
- A folder that holds several repositories opens the combined workspace review.
- Everything after that follows the chosen directory: staging, drafts, viewed marks, Guided Review and the feedback label.
Extra words still work as before.
/plannotator-review please check my changes, or evenlook at the api/users code, reviews the current repository and tells you which words it ignored. A mistyped path given on its own, such as./backnd, stops with an error instead of quietly opening the wrong repository. A PR link anywhere in the words now opens that PR.(#1644, closing #896 and #1482, by @oorestisime)
Additional Changes
- Keyboard scrolling works as soon as a review opens. Down, Page Down, Space, Home and End did nothing in plan review and annotate until you clicked the document. They now scroll it right away, without moving focus or interfering with shortcuts, dialogs, comment boxes or vim mode. Code review is a follow-up (#1648, closing #1647, reported by @henriquegeremia).
- No more Pi skill conflict warning. With both the install script and the Pi extension installed, Pi warned on every start that two
plannotatorskills collided. The extension now adds its copy only when no other one is loaded, so extension-only users still get the skill (#1643, closing #1642, reported by @tekumara). - OpenCode plugin on stable packages. The plugin now builds against the published
@opencode-ai/pluginfor OpenCode 1 and@opencode/pluginfor OpenCode 2, instead of a nightly build. The published plugin is unchanged in behavior (#1640, closing #1196, by @oorestisime). - OpenCode 2 checked on every change. CI now installs a current OpenCode 2 and loads the packed plugin into it (#1631).
- Link widgets for hosts.
@plannotator/uire-exports the editor's link-widget API, so apps that embed Plannotator's editor can draw their own chips for links (#1633).
Install / Update
macOS / Linux:
curl -fsSL https://plannotator.ai/install.sh | bashWindows:
irm https://plannotator.ai/install.ps1 | iexClaude Code Plugin: Run
/pluginin Claude Code, find plannotator , and click "Update now".Pi: Update
@plannotator/pi-extensionto 0.27.23 and restart Pi.OpenCode: Clear cache and restart:
rm -rf ~/.bun/install/cache/@plannotatorWhat's Changed
- ci(opencode): smoke the v2 plugin against a current OpenCode 2 by @backnotprop in #1631
- feat(review): persist viewed-file progress across review sessions by @oorestisime in #1632
- feat(ui): re-export linkWidgets for host link widgets by @backnotprop in #1633
- feat: opt-in background auto-update by @backnotprop in #1636
- feat(ui,core): in-document question blocks, core and ui by @backnotprop in #1637
- feat(editor): question answers, history, progress chip, Questions panel, Send answers by @backnotprop in #1638
- feat: question answers reach the agent by @backnotprop in #1639
- chore(opencode): migrate plugin dependencies to stable APIs by @oorestisime in #1640
- feat(review): Bitbucket Cloud PR review behind a PR provider seam by @backnotprop in #1641
- fix(pi): stop the plannotator skill colliding with the CLI install by @backnotprop in #1643
- feat(review): accept repository and worktree directory targets by @oorestisime in #1644
- fix(editor): scroll keys work on load without clicking the document by @backnotprop in #1648
- fix(review): path-shaped prose words no longer fail the review by @backnotprop in #1649
Contributors
@oorestisime contributed three of this release's pull requests. They built durable viewed-file progress for code review, moved the OpenCode plugin onto the stable OpenCode packages, and added directory targets to
plannotator review, each with thorough tests and clear notes on the tradeoffs.Community
- @drj613 asked for automatic updates in #1634, and @nclark added that they rarely update for exactly that reason
- @Toparvion explained how their team uses Plannotator with Bitbucket Cloud and Guided Review in #1583, and @bahayden first asked for other PR platforms in #671
- @dm3ch asked to review a specific worktree in #896, and @ben-reitz asked to pick a repository in a multi-repo workspace in #1482
- @tekumara reported the Pi skill conflict in #1642
- @henriquegeremia reported that keyboard scrolling needed a click first, with exact steps, in #1647
- @sergical opened #1196 to move the OpenCode 2 adapter to the stable plugin API
- @ashish921998 weighed in on keeping viewed state across review iterations in #1136
Full Changelog :
v0.27.22...v0.27.23 -
🔗 r/Harrogate HARROGATE EVENT THIS SATURDAY - 3RD OCTOBER rss
| Hey team! Me again :) Just another reminder of the new event running this Saturday, 3rd October, at Major toms:) 8pm - late Come along have a chat, a drink & a dance 🫶 Hope to see you there ❤️🩹🦎 submitted by /u/EllsMcL
[link] [comments]
---|--- -
🔗 earendil-works/pi v0.99.2 release
New Features
- MCP servers stay out of the way: servers with the default
codemodeexposure are no longer listed in thecodemodedescription and no longer block the first prompt. They appear in a short system prompt section, and scripts find their tools withsearchTools()anddescribeNamespace(). See Control tool exposure. - More MCP authentication options:
oauth.clientNamefor servers that only accept known OAuth clients, and"auth": { "provider": "<provider>" }to authenticate HTTP servers with a provider's/logintoken. See Authenticate with OAuth. - Anthropic workload identity federation from the Anthropic SDK environment variables. See Use an API key from the environment.
/reloadenables tools newly added to thedefaultToolssetting. See Tools.
Added
- Added a
descriptionfield for MCP servers (pi mcp add --description), shown with the server in the system prompt and used to rank its tools in tool search, and adescribeNamespace(name)codemode helper that returns a namespace's instructions and tool names.describeNamespace()andsearchTools()accept a namespace asmcp__dev-radius,mcp__dev_radius,dev-radius, ordev_radius. - Added an
oauth.clientNamesetting for MCP servers (pi mcp add --oauth-client-name) to change the client name sent during OAuth client registration, for servers that only accept known clients (#10226). - Added
"auth": { "provider": "<provider>" }for HTTP MCP servers to send a provider's current/logintoken as the bearer token instead of using MCP OAuth. The token is read on every request, so provider refreshes apply. Only allowed in the globalmcp.jsonand from extensions, and requires https except on loopback hosts. - Added Anthropic workload identity federation from the
ANTHROPIC_FEDERATION_RULE_ID,ANTHROPIC_ORGANIZATION_ID, andANTHROPIC_IDENTITY_TOKEN_FILEenvironment variables (see Providers) (#10177, #10242 by @philfreo). /reloadnow enables tools newly added to thedefaultToolssetting. Tools removed from it stay enabled, tools turned off during the session stay off unless newly added, and--tools,--no-tools, and--no-builtin-toolsstill override the setting (#10245).
Changed
- MCP servers with the default
codemodeexposure no longer appear in thecodemodedescription; scripts find them withsearchTools().codemode-deferredis now an alias forcodemode. Usedirectexposure for tools the model should see without searching (#10212). - The
codemodedescription no longer includes deferred tools, tool counts, or MCP server instructions, so it no longer changes when MCP servers connect or change their tools. Thetool_searchdescription no longer lists the servers whose tools it can load, for the same reason. Servers are listed instead in anmcp_serverssystem prompt section with a one-line summary, updated at the start of each prompt; a changed section is appended to the conversation. Scripts read server instructions withdescribeNamespace()(#10212). - The first prompt no longer waits for MCP servers without
directtools. They connect in the background and are waited for when a codemode script names them, a script searches tools, ortool_searchruns (#10212).
Fixed
- Fixed new sessions intermittently ignoring the saved default model, or warning that no models are available, when it belongs to an extension-registered native provider with a stored credential (#9962, #10190 by @davidbrai).
- Fixed the
/mcpsign-in URL not being clickable when it wraps across lines, by emitting it as a terminal hyperlink with aCmd/Ctrl+click to openline like/login(#10186). - Fixed codemode
image()accepting malformed base64 data or unsupported image types, which persisted an invalid image block that made every later provider request fail with HTTP 400 (#10215). - Fixed codemode failing to start its script worker from the standalone Windows executable (#10204).
- Fixed prompt submission slowing down with session length, because resolving the session's model selection looked up the model catalog once per assistant message (#10198).
- Fixed model lookups slowing down for providers with a refreshed pi.dev catalog, because merging remote catalog models took quadratic time.
- Fixed the
built-in-tool-renderer.tsandminimal-mode.tsextension examples removing the built-in tools' summaries and guidelines from the system prompt (#10072, #10193 by @christianklotz). - Fixed context overflow detection for Z.AI CN endpoint
Prompt exceeds max lengtherrors (#10208). - Fixed Anthropic requests failing when a tool schema uses keywords Anthropic strict tool use rejects, such as
minimum/maximum; such tools are now sent non-strict (#9953). - Fixed provider retries firing immediately when a
Retry-Afterheader contains an unparseable date; they now use exponential backoff (#9571). - Fixed extension commands registered without a string name or handler crashing pi when typing
/; the extension now fails to load with an error instead (#10054). - Fixed collapsed
codemodeand MCP tool results filling the screen when the output is one long line, such as minified JSON. Like bash output, the preview is now limited to wrapped lines instead of logical lines. - Fixed
codemode.mode: "only"listingread,bash,edit, andwritein the system prompt's tool list although requests only declarecodemode(#10192). - Fixed codemode scripts calling the wrong MCP tool when two tool names differ only in
-and_, such asread-fileandread_file. Like in Codex, MCP tool and namespace names now replace-with_(mcp__my-server__xis nowmcp__my_server__x), colliding tools of a server all get a hash suffix, and server names that differ only in-and_are rejected (#10239).
- MCP servers stay out of the way: servers with the default
-
🔗 anthropics/claude-code v2.1.286 release
What's changed
- Added a count such as "2 of 5" to the permission prompt when several permission requests stack up
- Added mouse support for the "N more" rows of lists in fullscreen mode: click one to jump to that end of the list, with hover and pressed states
- Fixed several Claude Code processes and IDE extensions each opening a login browser when gcpAuthRefresh or awsAuthRefresh credentials expire
- Fixed
claude --resumeand--continuesometimes losing every turn after a batch of parallel tool calls when the earlier session crashed or was killed - Fixed API 400 errors after a tool or hook returned an object, number or boolean instead of text, including in resumed sessions
- Fixed cloud sessions with very large histories never waking up because the container was stopped while the transcript was still loading
- Fixed the Claude apps gateway's spend meter pricing 1-hour prompt cache writes at the cheaper 5-minute rate, and counting only the first model call's input tokens on streamed turns that run a server-side tool such as web search
- Fixed macOS sessions still showing "Not logged in" or "Login expired" after
/loginsucceeds in another Claude Code window when a leftover~/.claude/.credentials.jsonexists - Fixed every turn failing when the Anthropic API refuses the model your default or a model alias resolves to: Claude Code now retries once on the previous model of the same tier
- Fixed Remote Control sessions (including
claude remote-control) staying connected after your organization's policy turns Remote Control off; they now disconnect with a notice - Fixed refusal and
--fallback-modelretries failing when the fallback model can't run fast; they now run at standard speed, with a one-time notice in interactive sessions - Fixed headless sessions repeating the "MCP servers require authentication" reminder after a successful re-authentication when the MCP discovery cache is enabled
- Fixed
claude auth statusreporting a Console sign-in's stored API key asclaude.ai; it now reportsapi_key, and the VS Code extension treats that session as an API key session - Fixed
/statuslisting an Anthropic profile beside an API key as if both were in effect; the profile is now marked as not in use - Fixed Claude not being told when a file attached to a message sent over Remote Control did not arrive, and a file sometimes getting only 10 seconds for its last download try
- Fixed a Remote Control message that arrived while Claude Code was exiting being marked delivered and then never answered; it now stays queued for the session's next run
- Fixed MCP error messages showing a credential's value when "Bearer" or "Basic" came before its key name
- Fixed percent-encoded Bearer tokens being only partly masked in error messages
- Fixed redacted logs and transcripts showing a secret whose key name has an invisible character inside, such as a zero-width space
- Fixed logs and transcripts showing part of a URL password that contains punctuation such as
), quotes,],&or a second@, or that runs past a/to a bracketed host such as[::1]in an ssh URL - Fixed the session transcript in the zip that
/feedbacksaves to disk containing invalid JSON lines after secret redaction - Fixed MCP connectors listing no tools for up to a day after their server dropped the older MCP handshake
- Fixed a repeat MCP sign-in request from Claude replacing the pending sign-in link, which could stop that link from working
- Fixed
/usagenot crediting an MCP server for a tool call made while that server was still connecting or had only just connected, such as right after a restart - Fixed plugins enabled on claude.ai occasionally going missing from Claude Code for a session after a transient server error
- Fixed a message typed into a running subagent showing twice in its transcript after the subagent read it
- Fixed subagent hand-back messages showing a raw task id instead of the agent's name when the subagent had no registered name
- Fixed foreground subagents sometimes missing the task-tracking tools (TaskCreate/Get/Update/List, TodoWrite) in sessions that have them enabled
- Fixed subagents spawned with worktree isolation loading the project CLAUDE.md and its imports a second time from the worktree copy on their first file read
- Fixed Workflow tool subagents being restarted from their original prompt when a connection stalled for a few minutes mid-response
- Fixed
/compact,/clear, and/rewindtyped while viewing a background agent's or teammate's transcript silently acting on the main conversation: a dialog now names the target and asks first - Fixed background jobs showing done while waiting for your approval
- Fixed the commit attribution reminder being re-sent inside tool output when a model fallback lasts only one turn
- Fixed a click on the space between words of a collapsed row (such as "Thought for 4s") highlighting the row without expanding it in fullscreen mode
- Fixed a row with no details, such as an action row with a long name, pushing every other row's details to the right in list screens
- Fixed files with very long names not reaching a cloud session when attached to it
- Fixed plugin errors for a marketplace Claude Code refuses to load: they now say why and how to fix it instead of "not found"
- Fixed
/plugin's Discover tab showing a marketplace name unquoted in its "Checking … for new plugins" line when its rows already show that name in quotes - Improved commit guidance: when your project or user skills include one named
verify, Claude is now told to run it right before committing, except for docs-only and tests-only commits - Improved send now (ctrl+enter) in a subagent's view: it now moves the subagent's running command to the background so your message is read right away
- Improved replies from background agents to your messages so they no longer open with a separate recap of what you said
- Improved claude.ai artifact link reads: WebFetch now asks the same questions as the Artifact tool's read (no artifact prompt while the session's network access is on, one per artifact while it is off), and an auto-mode yes no longer counts where only you can answer
- Improved fetch, skill, file read, sandbox network, Claude in Chrome, workflow script and notebook edit permission prompts to match the look of file edit prompts
- Improved Bash, PowerShell and Monitor permission prompts to show the command between dashed lines, matching file edit prompts
- Improved list scrollbars in fullscreen mode: in most lists the bar no longer shifts as the "N more" rows come and go, and it now has ↑/↓ arrows you can click or hold to scroll
- Improved the external editor (Ctrl+G): editors that take a line number now open on the line your cursor is on in the prompt
- Improved slash command suggestion responsiveness while typing when many skills or plugin commands are installed; command descriptions now match by word prefix
- Improved the output style picker: it now opens on your current style instead of Default, with each style's description on the line under its name; number keys no longer pick a style
- Improved
/hooks: the closing line of a hook's detail screen now says "this hook" instead of "it" - Improved the model fallback notice and the autocompact-thrashing error to say when a fallback dropped the context window from 1M to 200K tokens
- Improved responsiveness of SDK and
-psessions when a host re-sends an MCP server enable for a server that is already connected - Improved the protocol page a Claude apps gateway serves at
/protocol: it now says not to reject unknown input and matches what Claude Code sends today - Changed prompts sent while nothing is running or queued to show in the normal text color right away instead of gray
- Changed how failed API requests are retried: one limit now covers a whole model call, so with the default retry settings a failing call sends at most 14 requests
- Changed
--bareto connect only the MCP servers named on the command line, send the model no system reminders, and start no background tasks; under--bare, a shell command that reaches its timeout now stops instead of moving to the background - Changed the send-now key (ctrl+enter) to move a skill's own shell command to the background instead of ending it
- Changed the WebFetch error for a rate-limited domain safety check to tell Claude not to retry it in a loop
- Changed plugin installs to refuse npm sources that are git repositories or folders, and to install plugin dependencies only from registry packages
- Changed list screens (
/artifacts,/mcp,/skills,/hooksand others) to always line up each row's details in one column after the names - Changed the overflow rows of lists to read "↑ N more" / "↓ N more" instead of "N more above" / "N more below"
- Changed
/hooksto open on one list of your configured hooks grouped by event, so viewing a hook takes one Enter instead of three - Changed the theme picker to a scrolling list that fits your terminal instead of pushing the preview off screen; number keys no longer pick a theme
- Changed
/exit's Remove worktree to run after Claude Code stops the servers and shells it started there, which on Windows could keep the folder from being deleted - Changed the
claude-apiskill's Managed Agents examples to create environments with limited networking - Removed the browser link from
/ultrareviewandclaude ultrareviewoutput - Windows: Fixed
claude --bgand the agents view refusing a folder thatclaudealready trusts when its trust record was saved with different letter case - [VSCode] Added bookmarks: save Claude's responses and keep them in view in a Bookmarks side panel
- [VSCode] Added the questions Claude asks and your answers to the conversation: after you answer a question card, a Questions row shows each question with your picks
- [VSCode] Added option previews to question cards in the chat panel: the highlighted choice's mockup or snippet shows beside or under the options
- [VSCode] Added rows under a message that open to the terminal output, browser tab, browser instructions and selected code sent with it
- [VSCode] Fixed a second copy of a conversation opening in a tab when it was already open in the side bar; the side bar now switches to it
- [VSCode] Fixed settings dialogs reporting a failed save, without re-checking, when Claude Code printed more than 1 MB of output
- [VSCode] Fixed an endless "Teleporting session…" spinner when the extension stops responding
- [VSCode] Improved the Manage plugins dialog: it says when a turned-off plugin is still on because of other settings, and explains a plugin folder clash
- [VSCode] Changed Stop and Escape to end only the current turn; background agents keep running and can be stopped one by one from the agent map
- [VSCode] Changed the "✻ Claude Code" status bar item to show in every window, so you can open Claude when no file is open
- [Cloud sessions] Fixed an answered question card or approved tool call getting no reply when the session had gone idle after Claude sent a message
- [Cloud sessions] Fixed clearing an organization environment's setup script in admin settings leaving new cloud sessions still running the old script
- [Cloud sessions] Fixed the Runner actions menu on the self-hosted environments admin page closing on its own a few seconds after it opened
- [Cloud sessions] Fixed routine runs whose cloud session never started showing as Succeeded in the Runs pane, the routine's page and the sidebar; they now show as Failed
- [Cloud sessions] Fixed clicking an audio or video file in a cloud session's Outputs card opening an empty file search instead of playing the file
- [Cloud sessions] Changed a routine's page to read "Due" with the scheduled time, instead of a next run time in the past, when a scheduled run is late and hasn't started
- [Claude Tag] Added an Add channel button to Claude Tag's spend limits page in admin settings, so a limit can be set on any channel, including a private one, from its channel ID or Slack link
- [Claude Tag] Fixed memory recall finding nothing in organizations that cannot use the default Sonnet model
- [Claude Tag] Fixed a Slack channel losing its Claude settings when an Enterprise Grid admin moves it to another workspace and the first post afterward doesn't mention Claude
- [Claude Tag] Fixed Claude in a Slack thread occasionally starting over on a fresh machine, losing work it hadn't pushed, when your reply answered a question it had just asked
- [Claude Tag] Fixed public channel names on Claude Tag's spend limits page in admin settings showing as raw Slack IDs in larger organizations
- [Claude Tag] Improved the titles of sessions started from Slack as shown on claude.ai: they now read as the words you typed, without Slack user IDs or escape codes
-
🔗 OmniNull/OmniWM OmniWM v0.7.4 release
What's New Since 0.7.3
Name your windows, place the workspace bar on any screen edge, and choose OmniWM’s interface language. OmniWM 0.7.4 also adds Niri controls and moves blocking focus and window queries off the main thread.
Window marks and language support
- Find windows by name. Give live windows marks from the command palette or
omniwmctl. Mark matches appear first in window search; focus a marked window across workspaces or summon it beside the focused tiled window. Marks last until the window closes or OmniWM quits. - Assign mark shortcuts. Set Mark and Remove Mark are available in Settings > Hotkeys and start unassigned. The palette also provides buttons and shortcuts for marking its selected window.
- Bring recent windows forward. With an empty search, Windows mode lists the focused and recently focused windows first.
Shift + Entercan move the selected window into an empty current workspace, including floating windows; floating windows cannot be summoned to the right. - Choose the app language. Settings > General now lets OmniWM use a different interface language from macOS. The default follows macOS, and changes apply the next time OmniWM starts.
Workspace and layout behavior
- Place the workspace bar on any edge. Choose top, bottom, left, or right placement, with matching layout space, menus, and previews. Hidden Bar follows bar size changes and reveals concealed icons when the global workspace bar is disabled.
- Use shortcuts above workspace nine. Assign Switch, Move, and Move Column shortcuts to higher-numbered workspaces, including dynamic workspaces.
- Skip empty workspaces consistently. When Hide Empty Workspaces is enabled, relative navigation, workspace swipes, and Overview skip inactive empty workspaces. Direct workspace selection remains available.
- Control Niri sizing and gaps. Grow and shrink increments are now configurable from 1% to 100%, defaulting to 5% instead of the previous 10%. A new Gaps at Screen Edges option lets inner gaps appear only between windows.
- Keep Niri positioning predictable. Explicitly centered first and last columns stay centered during ordinary relayouts. A brief pill shows the column’s tabbed state, including when it contains only one window.
- Avoid Dwindle fullscreen overlap. Arriving tiled windows, including restored windows, end layout fullscreen before being added.
Responsiveness and feature controls
- Less blocking work during keyboard navigation. Niri keyboard and wheel focus now front apps off the main thread and target the chosen window directly, avoiding an initial activation of the app's previously active window. Layout passes skip a roughly 9 ms Screen Recording permission check once the wallpaper preview is captured, and Hidden Bar no longer rescans running apps on every app switch.
- Measured keyboard gains. In captures with matching display-link timing, estimated missed animation frames per focus-changing keypress fell from 2.5 to 0.125. Median hotkey handling fell from 6.5 ms to 0.44 ms. Median time to the first animation callback fell from 18.7 to 9.0 ms across apps and from 10.9 to 7.1 ms within the same app.
- Slow window queries move off the main thread. Ordinary frame-change WindowServer queries, including one recorded half-second wait, now run on a worker. New-window lookups and window lifecycle and rule metadata queries also run off the main thread. Closed windows release their cached previews.
- Choose which features run. Overview and workspace bar hover previews can be disabled independently. Disabled features release their panels, observers, previews, and retained resources; the global workspace bar switch now takes precedence over monitor overrides.
- Updated the embedded GhosttyKit terminal framework to
1.3.2-main+76895d97b(upstream changes). - Diagnostic captures provide more detail about focus, parking, frame reads, and preview activity.
Performance figures come from instrumented Release development captures on one Mac. They measure internal handling and animation callback timing; screen presentation latency was not measured.
Upgrade notes
settings.tomlupgrades to schema 4 on first launch. The original is kept assettings.toml.pre-v4(settings.toml.pre-v4.1if needed); the rewritten file removes comments. Restore the backup before downgrading to 0.7.3.- Direct IPC clients must use protocol 17. The
omniwmctlincluded in this release already uses it.
Thanks
Thank you to:
- binghan1227 for keeping the Hidden Bar icon aligned with workspace bar changes.
- Dereck Angeles for skipping empty workspaces during relative navigation.
- georgian for bottom and side workspace bar placement.
- Janos Miko for workspace shortcuts above nine and the Niri screen-edge gap option.
- Richard Ginzburg for preserving centered Niri edge columns, tabbed-mode feedback, and locale-independent formatting tests.
- Taylor Bell for runtime window marks and command palette improvements.
- Michael Künneke for supporting OmniWM through GitHub Sponsors.
Full changelog: v0.7.3…v0.7.4
Website and documentation · Installation guide
Release Integrity
The app is signed, notarized, and stapled. SHA-256 hashes:
OmniWM-v0.7.4.zip:1f1485e1d166831697e8cedb9c117577f5949fe00deeae8572bed6d6c58eb4e2GhosttyKit.xcframework-v0.7.4.zip:59e16cccd81ed87fde098bc40122154e640b950ba3b7b997879c786730b88c8e
- Find windows by name. Give live windows marks from the command palette or
-
🔗 crosspoint-reader/crosspoint-reader 1.6.5 release
Summary
Library view
Recent Books has grown into a powerful way to browse your entire collection on your SD card. Sort by recently added, title, or author, or just search for the book you want.
TTF support
Devices with external RAM (X4Pro, Sticky, X4C, PaperMono) can now use TrueType (.ttf) fonts directly — no need to convert them to CrossPoint's
.cpfontformat first. Drop your fonts into/fonts/or/.fonts/on the SD card, restart the reader, and pick them from the text settings. If a family has separate regular, bold , italic , and bold-italic files, put them together in one family folder..cpfontfonts still work everywhere, including on devices without external RAM. OpenType (.otf) fonts are also supported, though compatibility may vary.Cover Grid home theme
The new Cover Grid theme displays your recent books as a grid of covers on the home screen. This theme is not available on x3 and the original x4.
More control over reading
New word and character spacing controls let you adjust how tightly text sits on the page, alongside the existing line-spacing settings. Footnote navigation now highlights references directly on the reading page and and bookmarks can finally be renamed.
Touch controls and navigation
Touch devices gain configurable tap and swipe gestures for turning pages. Headers now have back buttons, and fixes improve scrolling in reader lists and settings. X4pro also gain configurable shortcuts for Home button.
X4 Classic support
The ESP32-S3-based X4 Classic is now officially supported. Get one now
Sleep Screens
Sleep screens across all devices now show drastically higher quality grays. On the X4 and certain Pro display variants they push the quality to the max. Transparent pngs got a nice quality bump as well.
Clock & About
You can now view the clock on the home screen! In addition, the old utc offset picker is gone and is replaced with auto DST and a list of cities to pick from when setting the time. This setting has moved to the more obvious system settings. You'll also find a new About option in the system settings as well which is useful for identifying display controllers during debugging.
Everything else
Files can be renamed straight from the file browser. Portuguese gets hyphenation support, Korean text justification stretches spaces between words rather than within them, and the keyboard gains an Arabic layout.
KOSync now sends more precise EPUB reading positions. Fixes also address EPUB lists, hidden content, chapter-position displays, and end-of-book navigation.
This release reduces font and EPUB memory pressure, fixes USB drive disconnection, improves web file-transfer safety, and prevents EPUBs that fail to render their first page from reopening automatically after wake.
What's Changed
- chore: update pioarduino to 55.03.311 by @serialx in #3397
- fix: Fixes KOSync memory checks and reduces memory pressure by @itsthisjustin in #3412
- fix: pack font manifest catalog into one arena by @fain182 in #3398
- docs(issue forms): Correct links to scope/roadmap by @cassidyjames in #3433
- fix(reader): synchronize end-of-book menu selection by @Daviex in #3418
- fix: render NFD Hangul filenames from macOS transfers by @serialx in #3036
- fix: reader's menu book chapter current position by @unnamedd in #3437
- fix: stabilize X3 EPUB anti-aliasing by @uxjulia in #3439
- fix(input): wake the idle poll on raw button contact so short presses register by @Techneaux in #3463
- feat: HTTP serve static with Cache-Control and ETag headers by @shirok1 in #2560
- chore: add direct download links for PR artifacts by @Uri-Tauber in #3389
- fix(webserver): normalize every user-supplied path and escape file names in the files page by @s0lness in #3353
- chore: Consolidates grayscale capability checks and enables absolute grayscale for supported screens by @itsthisjustin in #3478
- fix: don't display elements with hidden HTML attribute by @jjharpham in #3390
- fix(KOSync): compare mapped KOReader sync positions by @WhoTheHeck in #3111
- fix: update OTA to recognize the new format by @Uri-Tauber in #3493
- fix(debugging_monitor): if PSRAM is logged, add subplot by @olifre in #3490
- fix(KOSync): preserve precise KOSync upload progress positions by @WhoTheHeck in #3174
- refactor: reduce EPUB heap fragmentation with unique ownership by @serialx in #3518
- fix: reduce font-cache heap fragmentation by @serialx in #3521
- fix: release font caches before EPUB chapter layout by @serialx in #3527
- chore: add x4 Classic to CI pipelines by @Uri-Tauber in #3532
- docs: make roadmap easier to scan by @fain182 in #3517
- fix: number ordered lists and fix list container indents by @jan-xyz in #3500
- fix: dropped presses while a list repaints by @Techneaux in #3534
- perf: batch SdFat's SPI transfers on ESP32 by @osakanataro in #3501
- fix: USB OTG not disconnected when you unplug the cable by @itsthisjustin in #3538
- feat: Library view by @Uri-Tauber in #3366
- fix: Skip bw rendering on sleep images & fix white as transparent for sleep covers by @itsthisjustin in #3541
- chore: Resolve CI Node.js warning by @Uri-Tauber in #3545
- fix: Let OPDS catalog screens auto-sleep by @Uri-Tauber in #3547
- feat: Rename bookmarks by @Uri-Tauber in #3549
- fix: handle progressive JPEGs with separated component scans by @lpla in #2925
- fix: warp lists in library by @Uri-Tauber in #3556
- feat: add home button shortcuts by @uxjulia in #3516
- feat: Add AboutActivity for device information display by @itsthisjustin in #3563
- perf: keep settings list at its initial allocation by @serialx in #3569
- feat: Replace UTC offset with timezone and DST settings by @itsthisjustin in #3562
- perf: reduce image related heap fragmentation by @lpla in #2332
- fix: preserve frontlight state across silent restarts by @uxjulia in #3575
- feat(keyboard): add an Arabic layout by @rxmmah in #3430
- chore: Update license by @Uri-Tauber in #3578
- Swaps IMU direction on Sticky for proper gyro tilts by @itsthisjustin in #3583
- fix: clear ligature views when releasing SD font caches by @serialx in #3581
- fix: restore space widths on partial cache misses by @serialx in #3585
- fix(ci): pin pioarduino core inside the platform penv by @serialx in #3594
- feat: rename files by @Uri-Tauber in #3572
- fix: Re-Enable toolbar reader menu on button-only devices by @itsthisjustin in #3603
- fix: stop materializing every File Browser row (FUI list rowProvider) by @itsthisjustin in #3600
- fix: tear down WiFi when leaving the Settings network screen by @uxjulia in #3613
- feat: choose which UI languages are compiled in by @eszter007 in #3618
- chore: Update German translation by @BlueDragon333 in #3621
- fix: restore NFC composition in File Browser rows by @serialx in #3630
- feat: add word and character spacing controls by @serialx in #3528
- feat: refresh library by @Uri-Tauber in #3608
- fix: unblock X4 Pro frontlight double-click when short power press = Sleep, add Home key option by @naydichev in #3089
- feat: Adds Cover Grid UI Theme for PSRAM capable devices by @itsthisjustin in #3657
- feat: Add CrossInk style swipe and tap controls by @itsthisjustin in #3586
- perf: reduce per-pixel overhead in font rendering by @serialx in #3633
- feat: Convert slider dialogs to popups over current page by @itsthisjustin in #3669
- fix(settings): fix list scrolling by @brianhuster in #3668
- fix: Split navigation modes for cover grid & add selected states by @itsthisjustin in #3670
- fix: stop duplicating identical interval tables across font styles by @eszter007 in #3616
- fix: wrap OTA completion power-on hint by @lpla in #3632
- feat: Add TrueType fonts on PSRAM boards by @itsthisjustin in #3646
- fix: Sticky env missing ram stats by @Uri-Tauber in #3683
- fix: Reader list not scrollable with touch by @itsthisjustin in #3693
- feat: add Portuguese hyphenation support by @type0labs-dev in #3643
- chore: address Python Sonar warnings by @lpla in #2515
- feat: better footnotes navigation by @Uri-Tauber in #3682
- fix: Switch tap zones for RTL books by @Uri-Tauber in #3709
- fix: justify Korean text by word spaces only by @serialx in #3700
- feat: add back buttons to the header and unify all theme headers by @itsthisjustin in #3689
- fix: Remove loading icon display during startup by @itsthisjustin in #3671
- fix: release SD-font caches on reader exit by @serialx in #3699
- perf: avoid main-loop contention during EPUB rendering by @serialx in #3652
- fix: Toolbar menu in landscape mode + TTF font sizes by @Uri-Tauber in #3748
- fix: don't reopen unreadable EPUB on wake by @SurayaAtouraya in #3724
- fix: Remove article stripping logic by @Uri-Tauber in #3749
- chore: Update translations by @Uri-Tauber in #3396
- fix: don't cut screenshot folder names mid-letter by @SurayaAtouraya in #3750
- fix: speed up long status titles by @Uri-Tauber in #3573
New Contributors
- @cassidyjames made their first contribution in #3433
- @Daviex made their first contribution in #3418
- @unnamedd made their first contribution in #3437
- @Techneaux made their first contribution in #3463
- @shirok1 made their first contribution in #2560
- @s0lness made their first contribution in #3353
- @jjharpham made their first contribution in #3390
- @olifre made their first contribution in #3490
- @osakanataro made their first contribution in #3501
- @eszter007 made their first contribution in #3618
- @naydichev made their first contribution in #3089
- @type0labs-dev made their first contribution in #3643
- @SurayaAtouraya made their first contribution in #3724
Full Changelog :
1.6.0...1.6.5
Downloads
crosspoint-1.6.5-papermono.bin
crosspoint-1.6.5-sticky.bin
crosspoint-1.6.5-x3-x4.bin
crosspoint-1.6.5-x4c.bin
crosspoint-1.6.5-x4pro.bin -
🔗 exe.dev Show and Tell: Shelley rss
Shelley learned a new trick recently: instead of typing a prompt, you can record both your speech and your window (or screen). Once you stop recording, Shelley puts together the transcript and the matching video, and hands that off as the prompt to the agent.

You have to see this to believe it. Shelley had a bug where if you sent off a “Compact and Send” and navigated to a different conversation, the navigation would revert back to the original one. I could have described it, but after it happened to me, I started a new conversation, and took the following recording:
Your browser does not support the video tag.
The recording is not edited except my shortcuts and URL bar are cut off, some commentary is added below, and it’s sped up. It’s painful to watch (especially for me!): all of those ums, and figuring out what to click to reproduce, and so on. And yet, the next message from me in that conversation was “push.” The agent was able to reproduce based on what it saw and fix the bug, in one-shot, and I typed nothing.
Give it a ~~shot~~ screen recording!
-
🔗 mhx/dwarfs dwarfs-0.15.8 release
Bug fixes
-
Correctly handle sparse files when querying entries in "data order". This is primarily used by
dwarfsextractto speed up extraction and avoid repeatedly decompressing the same blocks. All regular files are sorted by the block referenced by the first chunk. While subsequent chunks may also reference earlier blocks, these blocks are typically still in the cache. However, a sparse file starting with a hole will reference a special "hole marker" block in its first chunk, causing all sparse files that start with a hole to be clustered at the end of the list, without considering any of the subsequent data chunks. This is likely an academic edge case, but in an image containing tens of thousands of such sparse files, this can quite dramatically slow down extraction. Furthermore, this also triggers other bugs that can quite easily causedwarfsextractto run out of memory. This has been fixed by ignoring the holes and only considering the first data chunk when sorting. -
Limit size of worker group queues in filesystem extractor. During extraction, the filesystem extractor used by
dwarfsextractuses a worker pool for writing regular files. This pool has a job queue that by default can grow unbounded. Each job, before being added to the queue, will already trigger an asynchronous read of the file data. This is actually bounded based on the size of the queued regular files, but not considering the amount of memory used by the decompressed blocks in order to actually read the queued file data. This usually isn't a problem, since the files are added in "data order", but in combination with the bug described earlier where sparse files starting with holes end up not actually being sorted, this can cause requests for pretty much all blocks in the image to be decompressed simultaneously. Even worse: if the decompressed blocks are larger than the configured cache size, blocks that are still referenced by eariler requests can be evicted from the cache, causing multiple simultaneous decompression requests for the same block. So even though the decompressed image might actually fit in memory, the actual memory consumption can very easily reach tens of gigabytes. Limiting the queue size is a workaround for this issue, but it's a good idea to bound the queue size anyway. There are two real fixes: one is the fix for the sorting issue described earlier, the other one is for the block cache to keep track of blocks that are evicted, but still in use, and re-insert these blocks into the cache rather than decompressing them again. The latter is a more involved change, so it is not included in this release, and it's really not necessary for thedwarfsextractuse case with the former fix in place. -
Correctly compute total hardlink size. The algorithm used to compute the total hardlink size (i.e. the total size of all hardlinks beyond the first one) was not correctly accounting for non-hardlinked duplicates. The total hardlink size is currently really only used for progress reporting in
dwarfsextract. However, v0.15.0 introduced a strict check that would abort the process if there was a mismatch between the total computed size and the total written size if--stdout-progresswas used. This strict check would fail if the total hardlink size value was actually used, which is only the case for extracting to formats that do not support hardlinks, e.g. ZIP or 7z. A workaround would be to simply drop--stdout-progress, or to make the strict check just report an error rather than aborting (this is also done in this release). The real fix is to correctly compute the total hardlink size, and fix a wrong value when re-writing an image and using--rebuild-metadata. -
Don't abort extraction in release builds when discovering a progress mismatch with
--stdout-progress. The mismatch is still reported as an error, but the process will continue in a release build and abort in a debug build. -
Preserve preferred path separator when rebuilding metadata. When rebuilding metadata, the preferred path separator should never be changed, but it was erroneously always set to the platform default of the binary used for the rebuild. So an image created on Windows with
\as the preferred path separator and later rebuilt on Linux would end up with/as the preferred path separator (and vice versa). This would only cause problems if the image contained symbolic links with target paths that used a path separator. Those symbolic links would now be broken. A workaround would be to rebuild the image on the same platform it was created on, with a version of DwarFS that still contains the bug (i.e. v0.15.7). This is now fixed and the preferred path separator is preserved when rebuilding metadata. -
Call
archive_write_finish_entry()to ensure that warnings are reported for all entries indwarfsextract. Before this fix, all warnings except for the last were silently ignored. -
Handle sparse files when using readahead. The readahead implementation did not handle hole chunks in sparse files and would actually crash in debug builds. This has been fixed.
-
Default
max_eager_map_sizeto 1 TiB on 64-bit architectures. This defaulted tounlimitedon 64-bit architectures before, which could cause errors when attempting to map extremely large (sparse) files.
Build fixes
- Add several missing includes to avoid compilation errors on newer compiler versions.
Full Changelog :
v0.15.7...v0.15.8SHA-256 Checksums
ee100574a1340ad054bdf7543bedde3bd040639114ede65f1c20e08567eff2b1 dwarfs-0.15.8-Linux-aarch64.tar.xz 484727f2efdf44f597abe4950d780b9fcfc0a499672338dc1ba99099d92f252e dwarfs-0.15.8-Linux-arm.tar.xz 71298018cdd73f5de267b4af33c015df840adf18c6131d8ca7b162dc8e20226b dwarfs-0.15.8-Linux-i386.tar.xz 974056bafd72731f0b18f324211e903031cb3809e2c50ff505f22efd2f4af6c6 dwarfs-0.15.8-Linux-loongarch64.tar.xz 70e83eaff3dd28f966502c890971a79f23d69a7a0f19ca4d0f37791a888cba7e dwarfs-0.15.8-Linux-ppc64le.tar.xz 2c681ce53a9f29c0259f22700259fca2dcd4ff5ef8bd2715a01e210232e558b6 dwarfs-0.15.8-Linux-ppc64.tar.xz 78e1a66e1f3bd3da6967ba7590ae8ac30915c419c8d5fda75b19b08d570b51ac dwarfs-0.15.8-Linux-riscv64.tar.xz da0b5e6995b2a7679cb950610fc00dd3bad663623e0a55b49b5f342e8a20d5fd dwarfs-0.15.8-Linux-s390x.tar.xz f5eb34c3e2ac1bfcbe66e15b8a66da20897049b22867d2c193bffedf25cc08ae dwarfs-0.15.8-Linux-x86_64.tar.xz a2382a2d06f4539b1c53b8b4f800776945e2f13f71c0a3226d3bab3e1b25fe04 dwarfs-0.15.8.tar.xz 49ffae19533b0c5b3dfc5516b19e97700f1197416062a5517f65bfb631631566 dwarfs-0.15.8-Windows-AMD64.7z e6babb706254fd98b7012325c78997983a7824c097f87706efc0fbd365c04fd8 dwarfs-fuse-extract-0.15.8-Linux-aarch64 3b233d30051ecfe4cd4d9bb33d0b9001d99c96f0e4067a3d5d5d7b1ea71733e0 dwarfs-fuse-extract-0.15.8-Linux-aarch64.upx 789f07f790e28ee198b5854364759b579b3744407d22f17bb2170e7a2c006bb6 dwarfs-fuse-extract-0.15.8-Linux-arm caf35bda41970c426e1fa162bf7340eef80001a7ec2cb59643e04b3ab0d89a00 dwarfs-fuse-extract-0.15.8-Linux-arm.upx ec796f13b5757ce516d4d4b7a172aebe5aa6cded7a4388e8e9eb873a8291bf98 dwarfs-fuse-extract-0.15.8-Linux-i386 594b838742d9310448adaa21fafd084c46e7e0bf319254045d31374fe257466d dwarfs-fuse-extract-0.15.8-Linux-i386.upx 6fccfa6437467eaa6ee59051bd7148989b3b5eaeefbbeedbb092b72d9808721a dwarfs-fuse-extract-0.15.8-Linux-loongarch64 5709f9fe9f5d492969e860af226777ece0bc88a4d61bcf2fd5f919d1c433e786 dwarfs-fuse-extract-0.15.8-Linux-ppc64 cd6832bccb37f58312621f4c6c1b72f623d667a366867811cca8451a2524f18a dwarfs-fuse-extract-0.15.8-Linux-ppc64le 94fc718d54b7b1164df5405e9e319043d263cfc62bad7f56b4c9031d75748779 dwarfs-fuse-extract-0.15.8-Linux-riscv64 13bbfc0fa36d3527b56726d480119a4b920c372d4f0a47b8cc01d492a951de40 dwarfs-fuse-extract-0.15.8-Linux-riscv64.upx 3724678c157108c10f1c6560923ef1319a24011a94744c6a6ccae3ac03d0298b dwarfs-fuse-extract-0.15.8-Linux-s390x 494030845b6b61efe74bd00449069faede1720ef5b8e0f59ccc5eea7d74a244e dwarfs-fuse-extract-0.15.8-Linux-x86_64 0cbe87b4b6961a1003d3a4dc704a0271148551a4584405857d6bff1b6aa921d7 dwarfs-fuse-extract-0.15.8-Linux-x86_64.upx f21cbc33840bbffe410c7043fab85e0d1b4eee86d77bcc0fbf3a51b94fc4b222 dwarfs-universal-0.15.8-Linux-aarch64 84330541b1d989637ddcc4b521bc33e2654e7bf962f7e488884cc87ff86081b4 dwarfs-universal-0.15.8-Linux-aarch64.upx c098d0a222e737a71764343e37c700f502e75c3de7fb1e4802410976f7053a23 dwarfs-universal-0.15.8-Linux-arm 9d1fdc9ad57166efdb52a900b6a3b8f7acae649e3de33c3b3f0b62d1a056b379 dwarfs-universal-0.15.8-Linux-arm.upx 5c93f65020ad20c0d0be8271d30d0a8da4bbfdec0a1078e6f0606658e7e110f6 dwarfs-universal-0.15.8-Linux-i386 a5f47ae2bd0a35340a5db9ce4b081a8fbcf5facff187bed869420d1a7ca73d01 dwarfs-universal-0.15.8-Linux-i386.upx e966c33a66b0ee00e4bf0705426c7261c4798018b9c67d719c833bdd48f1beea dwarfs-universal-0.15.8-Linux-loongarch64 0b45d8f8c09e5de71f9ea3d290871cfd3aef4aaac1aa92d8363b84e6c8f96c3a dwarfs-universal-0.15.8-Linux-ppc64 257fcab374891dadf5b5f9947997b9ad34951942f4f724dc18917411add9957c dwarfs-universal-0.15.8-Linux-ppc64le 2370e78a69f9255cd9ae0ca54c987f3dd706608e6d142972379f1a924de9a9e5 dwarfs-universal-0.15.8-Linux-riscv64 1cc6182f040cd7037563f57827cc0fdacf21df318e41d1e8f105568e95fa10b8 dwarfs-universal-0.15.8-Linux-riscv64.upx baa138983599c8a8ecd3266c77c4c8794b96e587d1bef667ecef53bd4dea3a3a dwarfs-universal-0.15.8-Linux-s390x 0ed66df82f9bb6671c1f898404ff4d57c263460078d46dbac8efb9d8eec271b7 dwarfs-universal-0.15.8-Linux-x86_64 f149494612fa4d6b2a3b49892af25f629ed0e241da1686924a0f4f121f1ff521 dwarfs-universal-0.15.8-Linux-x86_64.upx 7bce40f0652c52c9444d8029af5821d77e38a9360d9dfab52c4dbecd6180506d dwarfs-universal-0.15.8-Windows-AMD64.exe e78d20df86b1382b44e0e15b3b7fc69402878ed8552eb218544f0336267a69ec dwarfs-universal-small-0.15.8-Linux-aarch64 da540fb3961bb62f9e676ada3570c024501b87af63d5ac98e1074ce918cec53f dwarfs-universal-small-0.15.8-Linux-aarch64.upx 69846272903ff28ea3878146b9a323b37bc8d6addaa75372631eee1668c805d1 dwarfs-universal-small-0.15.8-Linux-arm 1e2648736308d66dec871f1ee970cb5c2c46419345e7da9b182df5d8166cf9f2 dwarfs-universal-small-0.15.8-Linux-arm.upx 28a8027220c084e82ae2d6f78789edbbc23d585f6a9ade86b326799719c06592 dwarfs-universal-small-0.15.8-Linux-i386 87f585381507387f9bfa230a7ff08acee7e5ef0eff52d3de920043ac98c04bdf dwarfs-universal-small-0.15.8-Linux-i386.upx 2316584997ee6441b483d859b184a986db92639efc5de899ad48c4e69c86a8c4 dwarfs-universal-small-0.15.8-Linux-loongarch64 fff8c7671a4cba109c48cf40aa919e0809bc59feda140e3685afd3a46b673f53 dwarfs-universal-small-0.15.8-Linux-ppc64 22e61cb0021f1bfba953f86b520f4f192e9b33163453a7074ba6dbf6d0f9bf81 dwarfs-universal-small-0.15.8-Linux-ppc64le 4ce7177f04f1e2ce149cadf461e098489f7e4fa0559ed0466e60b7c9320cd2e7 dwarfs-universal-small-0.15.8-Linux-riscv64 f58d0991f0c4ba5475ed4457df2ab08a28f261e0ed23b26277017dab8a48e059 dwarfs-universal-small-0.15.8-Linux-riscv64.upx 72489cfaf6e2a66992590a6075fc6681122c367b55c311d2bbed15dad590375a dwarfs-universal-small-0.15.8-Linux-s390x d29734ed51547f5a5e1c92ede3ea0f2e003902f5c58df10d5e060c87b4e07e69 dwarfs-universal-small-0.15.8-Linux-x86_64 0217b8bf803b934aa1fd097b1c131cff0ffc112a03f575996d93a7882caeb8d4 dwarfs-universal-small-0.15.8-Linux-x86_64.upx -
-
🔗 Cryptography & Security Newsletter Get Ready for Seven-Day Certificates rss
With the groundwork for the post-quantum migration of key establishment behind us, browser vendors are increasing their efforts on the remaining parts: post-quantum authentication. (If you haven’t been following the events surrounding post-quantum migration, our newsletter from two months ago has a quick recap. Read it before continuing here.)
-
🔗 smol-machines/smolvm smolvm v1.22.0 release
What's Changed
- Seed image machines on macOS and in the SDKs by @BinSquare in #1486
- Serialize the tests that depend on SMOLVM_MACHINE_NAME by @BABTUNA in #1490
- Seed ephemeral machine run from the shared image seed by @BABTUNA in #1491
- Let a machine stop itself when its workload exits by @BinSquare in #1444
- Cache recent checkpoint restores and add checkpoint-warm by @BinSquare in #1492
- Bump the workspace to 1.22.0 by @BinSquare in #1487
Full Changelog :
v1.21.1...v1.22.0 -
🔗 smol-machines/smolvm smolvm v1.21.1 release
What's Changed
- Stream macOS checkpoints sparsely so saves scale with memory in use by @BinSquare in #1484
- Reword the README description, add an SDK example, and explain shared responsibility by @BinSquare in #1483
- Bump the workspace to 1.21.1 by @BinSquare in #1485
Full Changelog :
v1.21.0...v1.21.1
-
- September 29, 2026
-
🔗 smol-machines/smolvm smolvm v1.21.0 release
What's Changed
- Let machine update change a stopped machine's egress allow list by @BinSquare in #1471
- Bump the Nix package to 1.20.2 by @BinSquare in #1472
- Support Windows memory checkpoints and frozen branches by @BinSquare in #1474
- Restore a checkpoint's layered disk as a copy-on-write top instead of copying its top layer by @BinSquare in #1475
- Let a stopping server finish in-flight requests for a configurable grace by @BinSquare in #1473
- Cut machine exec latency: wake on exit, stop crun re-cloning itself, fork less in the wrapper by @LoganGrasby in #1479
- Speed up cold checkpoint restore by hashing during extraction by @BinSquare in #1480
- Let embedders create detached machines and list every machine by @BinSquare in #1478
- Key a restored machine's VMM uid on its own data dir by @BinSquare in #1476
- Treat an empty machine start body as no body by @BinSquare in #1481
- Start registry-image machines on a shared seed of their image by @LoganGrasby in #1477
- Bump the workspace to 1.21.0 by @BinSquare in #1482
Full Changelog :
v1.20.2...v1.21.0 -
🔗 anthropics/claude-code v2.1.285 release
What's changed
- Added
CLAUDE_CODE_DISABLE_WEB_FETCHenvironment variable to turn off the WebFetch tool - Added
claude --desktopto open the Claude desktop app on the current directory, or on a session with--continue/--resume <id> - Added
claude plugin configure <plugin>to show a plugin's options and which are unset, or save new values read from stdin with--values-stdin - Added
<server>.<key>=<value>toclaude plugin install --config, so a bundled.mcpbMCP server's own settings can be set at install time and it starts without visiting/plugin→ Configure - Added
allowedProvidersmanaged setting to limit which API providers a machine may use (Anthropic API, a custom endpoint, Bedrock, Mantle, Vertex AI, Foundry, Claude Platform on AWS, or a Cloud gateway) - Added
CLAUDE_CODE_NONSTREAMING_TIMEOUT_RETRIESenvironment variable to cap re-sends of a non-streaming fallback request that timed out - Fixed
claude -pwithCLAUDE_CODE_FORK_SUBAGENT=1: a subagent's own Agent call now runs in the foreground, so the subagent gets the child's result - Fixed plugin and marketplace installs and updates over SSH ignoring the ssh program set in
GIT_SSHor in your git config'score.sshCommand - Fixed Claude Code refusing to start when the OS denies reading the managed settings file; it now warns and starts without that file's policies. Other read errors and unparseable files stop every session
- Fixed cloud sessions that restarted after their conversation was compacted refusing the next update to an artifact the session had already read or published
- Fixed
claude plugin disableandenablewith a fullname@marketplaceid changing a settings entry in another letter case instead of the installed plugin's own - Fixed files attached to a message sent over Remote Control being left out after a single failed download; a network error, timeout or server error is now retried up to twice
- Fixed switching models mid-session with a
set_modelrequest (such as the Agent SDK'ssetModel) leaving the new model on the built-in output-token limit and auto-compact window until restart - Fixed redacted logs and transcripts showing part of a URL password that contains
@, or all of it when the URL writes its@as%40 - Fixed SSH passphrase and new-host prompts from worktree and
/teleportfetches taking over the terminal; these fetches now fail fast instead of asking - Fixed switching off an MCP server added mid-session in SDK and
-psessions leaving its tools available - Fixed
claude -p --permission-prompt-tool: a background subagent's permission request now goes to the prompt tool instead of being auto-denied - Fixed
claude mcp listandclaude mcp get, and the not-found error ofclaude mcp remove,loginandlogout, printing line breaks and terminal escape sequences from MCP server names and values - Fixed sandbox auto-allow asking for approval on every run of many inline scripts (
python3 -c,node -e) just because they contain= - Fixed fork subagents not keeping the session's plan mode or
dontAskmode: a fork now runs under its parent's permission mode and cannot exit plan mode - Fixed
claude remote-control --helpsaying--[no-]chromedefaults to the machine's/chromesetting; spawned sessions keep Claude in Chrome off unless--chromeis passed - Fixed background subagents in auto mode prompting a second, redundant reply after each report
- Fixed cloud session creation and
/remote-envreading only the newest 20 of an account's environments - Fixed Remote Control marking a message as read as soon as it arrived instead of when Claude started on it, and losing a message still queued when the terminal quit (it now arrives on the next resume)
- Fixed installing a plugin with
claude plugin installor/pluginputting it into an installed plugin's cache or data folder when their ids differ only in.,-,@or (macOS, Windows) capitals; the install is now refused - Fixed hooks and SDK permission callbacks seeing a missing or outdated plan on ExitPlanMode when the plan was written in the same response
- Fixed the first reply in cloud sessions arriving tens of milliseconds late, a regression in 2.1.283
- Fixed sessions that authenticate with
ANTHROPIC_AUTH_TOKENagainst the Anthropic API never loading the organization's policy - Fixed a failed
agent(),parallel()orpipeline()call that a workflow script awaits later, or not at all, being treated as an unhandled promise rejection, which could end a background session - Fixed synchronous hooks hanging Claude Code while a background process the hook started (for example
some-daemon &) kept its output open; the hook now finishes shortly after its own process exits - Fixed WebFetch reporting a rate-limited domain safety check as a network or enterprise policy block
- Fixed the fullscreen ctrl+o transcript freezing briefly when opened on turns with hundreds of file reads or searches; tool calls still running when the transcript opens now show their results when they finish
- Fixed Amazon Bedrock mid-stream
modelTimeoutExceptionandserviceUnavailableExceptionerrors showing a raw JSON body instead of the error message - Fixed
/autofix-prand/schedulesaying the Claude GitHub App is not installed on a repository whose install status had not been checked yet - Fixed dismissing a row (x) in
/artifactsunlinking its file from the artifact, so publishing the same file again created a new artifact instead of updating it - Fixed Artifact tool publishes after a conversation rewind (Esc Esc) overwriting a file's newer content that Claude had read only in the rewound turns; the publish is now refused until Claude re-reads the file
- Fixed an
Artifactallow rule ("don't ask again") letting the Artifact tool publish a file outside the working directories without asking; add the file's folder with--add-dirfor the rule to cover it - Fixed
/costand SDKmodelUsagereporting a turn under the wrong model when the server answered a refusal with a different fallback model than the client expected - Fixed the Artifact tool so that publishing a page no longer lets Claude overwrite its source file without re-reading it when Claude's earlier read was cut short or the file had changed since
- Fixed auto mode skipping its classifier for Artifact tool asset uploads and reads of someone else's artifact when you had approved that artifact earlier in another permission mode
- Fixed a misleading "core.worktree is set" error from
/ultrareviewwhen the project folder briefly could not be read - Fixed
/ultrareviewon macOS and Linux failing to upload the working tree from a git worktree whose per-worktree config setscore.longpaths - Fixed the Artifact tool sometimes reporting a publish as a conflict with another session after retrying a temporary server error, when the first attempt had actually succeeded
- Fixed
/ultrareviewuploads including uncommitted changes to credential files whose name has a colon before the extension, such asserver:8443.key - Fixed a rare auth failure when two sessions recover a login refresh lock left by a crashed process at the same time
- Fixed the PowerShell tool's permission check skipping deny and ask rules, and caching that failure for later checks, when its command parser failed to start (for example when the machine was out of memory)
- Fixed
/ultrareviewuploads on macOS and Linux running slowly on some unusual file names, and their credential-file check missing file or folder names with many backup or editor marks - Fixed a cancelled shell command or hook still starting, and running to its end, when the cancel arrived while it was being set up
- Fixed vim mode: after editing in the external editor (Ctrl+G),
xorrin NORMAL mode no longer breaks a pasted-text placeholder at the end of the prompt - Fixed responses blocked by the API's output content filter being re-sent and retried, sometimes for minutes, instead of showing the filter's error right away
- Fixed plugins silently skipping a bundled
.mcpbMCP server that still needs configuration:/plugin, the install message andclaude plugin installnow say so and point to Configure - Fixed compacting or resuming a session failing, opening without its history, or crashing when its saved transcript holds a compaction marker or loop wakeup entry with missing or malformed fields
- WSL: Fixed
/ultrareviewrefusing to upload a checkout on a Linux volume when a changed file's name has a colon or ends in a dot or space - Fixed
CLAUDE_CODE_RESUME_INTERRUPTED_TURNre-running a turn that had ended at--max-turns - Fixed sign-in that could wait forever after the browser showed success
- Windows: Fixed
/ultrareviewuploading a linked worktree of a repository rooted at your home folder in some cases - Fixed cloud sessions reporting the uploads folder as missing before any file had been uploaded
- Fixed a reply sent from
claude agentsto a background session waiting on a permission prompt sometimes approving the pending command - Fixed
claude attach,logs,stop,respawnandrmstarting a new session with the command name as its prompt when options came before it, such as from a shell alias - Fixed
claude mcp listleaving out WebSocket (ws) MCP servers; each is now listed with its URL and health status - Fixed
claude mcp getshowing no Type, Command, Args, or Environment for stdio servers whose config entry omits thetypefield - Fixed
.claude/settings.local.jsonallow rules being held back outside a git repository when git's trace2 output is configured - Fixed the
/claude-apieval runner scaffold and report builder writing through a symlink or hard link planted at an output file - Fixed the
/claude-apieval runner scaffold counting responses cut off atmax_tokensin the score averages; they are now marked truncated and counted separately - Fixed
showing as literal text in the terminal when a reply uses it to indent text, such as row labels in a markdown table - Fixed a brief freeze (up to a second) partway through long sessions outside fullscreen mode, which came back after
/clearor/compact - Fixed an approved Edit never going through when its target is a device, such as a file symlinked to /dev/null, and the approval came from the IDE diff view or changed the edit
- Fixed a failing API request being retried up to 21 times when streaming kept failing; the non-streaming fallback now shares the request's retry budget instead of getting a fresh set of retries
- Improved Claude in Chrome: the native host now reports your computer's name, so connected browsers can be labeled by computer instead of "Browser 1" / "Browser 2"
- Improved Bedrock and Vertex AI sessions to switch to an older available model of the same tier, instead of failing, when an admin removes access to the default model; session titles and summaries now fall back with it
- Improved plugin marketplace errors to name why a git address is refused instead of citing enterprise policy
- Improved validation of git URLs for plugins, marketplaces and the current repository's remote
- Improved Artifact tool results: they now suggest publishing in the same step as writing or editing the page, which can save a round trip
- Improved Remote Control: a
/btwside question asked of a session hosted by an app such as Claude Desktop now sees the turn in progress, not only the last finished one - Improved pictures Claude sends as BMP, HEIC, HEIF, AVIF or TIFF files: the Claude apps now show a preview where Claude Code can convert them
- Improved subagents in auto mode: a subagent's run now ends as soon as it hands its report back to its caller, instead of taking extra turns that reach no one
- Improved Bedrock and Vertex start-up model checks: models your account cannot use are now remembered for up to a day instead of being re-checked on every launch
- Improved Artifact tool publish results to use fewer tokens: the note on updating an artifact is shorter, and where to find your artifacts is no longer repeated after every publish
- Improved SDK liveness during a non-streaming fallback request: with partial messages on, a
pingstream event is now sent every 30 seconds on the Anthropic API, Claude Platform on AWS and gateways - Improved
/resumeandclaude --resumeon a session that is running in the background: they now open that session instead of refusing, and a prompt given withclaude --resume <id> "prompt"is sent to it as its next turn - Improved per-turn performance when many permission deny rules and MCP tools are configured
- Improved responsiveness when leaving the ctrl+o transcript view in long sessions when fullscreen rendering is off
- Improved Bedrock, Vertex and Mantle start-up model checks to send the same User-Agent, x-app and session ID headers as regular requests
- Changed MCP tools so a tool that sets its own
_meta['anthropic/alwaysLoad']to false stays deferred when its--mcp-config, Agent SDK or plugin server is set toalwaysLoad - Changed background Bash and PowerShell commands to stop after a time limit (their
timeoutwithrun_in_background, default 30 min, max 2 h); Claude is notified when one is stopped - Changed Code Review's pull request reviews and
/ultrareviewto run whendisableWorkflowsis on, unless the machine running the review has it set by its own administrator (MDM or the managed-settings file) - Changed sessions behind a custom
ANTHROPIC_BASE_URLto use the 1M context window of models that have one (Opus 4.7+, Sonnet 5+, Fable); run/autocompact 200kif your gateway stops at 200K - Changed Team and Enterprise sessions, and sessions whose sign-in plan Claude Code can't determine, to withhold WebFetch until the organization policy loads if it couldn't be loaded at startup
- Changed /memory so that Auto-memory can no longer be turned on from a background session or from a session one of Claude Code's own tools started; turning it off there still works
- Changed the one-time offer to make auto mode your default permission mode to also show on third-party providers and with telemetry off, when your user settings default to another mode
- Changed
claude -pand Python Agent SDK sessions on third-party providers or with telemetry off to start in auto mode when no permission mode is configured, like interactive sessions;--permission-modestill overrides it - Changed Bedrock, Mantle and Claude Platform on AWS requests to a base URL with a non-default port to include the port in the SigV4-signed Host header
- Changed the MCP server name
widgetsto be reserved in cloud sessions and on self-hosted runners: your own server under it, or a close spelling such aswidgets_, no longer loads, so rename it - Changed
/ultrareviewon macOS and Linux to leave symbolic refs out when uploading a local checkout; a checkout whose current branch is a symbolic ref is now refused with an explanation - Windows: Changed project and local settings
envto no longer setALLUSERSPROFILE,SystemDrive, or theCommonProgramFilesvariables; set them in user or managed settings instead - Changed
/tasksto fold background work Claude Code runs for itself under one "System tasks" row; press Enter on it to show those tasks - Changed
/ultrareviewon macOS and Linux to require git 2.31 or newer to upload a local repository; checkouts made with--separate-git-dirare now refused instead of being uploaded with an older method - Changed
/ultrareviewuploads on macOS and Linux to send a partial clone as a working-tree snapshot on git 2.31 or newer, instead of falling back or refusing when git's version looked too old - Changed
/ultrareviewuploads on macOS and Linux to refuse, instead of fetching, a partial clone missing some of its working tree's files on older git versions; a clone made without--filteruploads - Changed Bedrock, Vertex and Mantle start-up model checks to identify themselves as Claude Code, like other Claude Code requests
- Changed
claude mcp getto hide the command, arguments, and environment values of stdio MCP servers provided by plugins; variable names are still shown - Changed
/claude-apiso it can no longer be run from Remote Control clients - Changed
/config chrome=trueto direct you to the /config panel instead of enabling Claude in Chrome by default;/config chrome=falsestill turns it off when it was on - Changed sandbox settings so project settings cannot widen or turn off an admin-required sandbox, replace the proxy behind a managed deny list, extend a strict allowlist, or reopen managed read-denies
- [VSCode] Added a note under a restored tab's last message when a window reload interrupted it and no reply will follow
- [VSCode] Added a plugin options form to Manage plugins: installing a plugin that has options asks for the unset ones, and a gear on its row changes them later
- [VSCode] Added an on-demand diagnostics tool so Claude in the panel can read the Problems panel's current errors and warnings at any time, not only right after it edits a file
- [VSCode] Fixed pressing Enter after typing a slash command running an unrelated menu item picked by fuzzy match, or doing nothing
- [VSCode] Fixed an open agent transcript losing the agent's newer messages during a long session
- [VSCode] Fixed a message that quotes a Claude Code or IDE tag losing the rest of its text in the chat
- [VSCode] Fixed a message sent while Claude was working disappearing from the conversation after the session was reopened
- [VSCode] Fixed the session list's Web tab showing the previous account's sessions after an account switch
- [VSCode] Fixed restored tabs re-running an interrupted turn when VS Code was started with
CLAUDE_CODE_RESUME_INTERRUPTED_TURNset, even with Continue After Reload off - [VSCode] Fixed Escape stopping the running turn instead of closing the command menu after clicking one of its rows
- [VSCode] Fixed opening Past conversations replacing a live conversation with its saved copy
- [VSCode] Fixed a Claude tab reloaded after an extension restart staying blank instead of saying how to recover
- [VSCode] Fixed opening a conversation that is already open in another window or app starting a second copy of it without warning; it now asks first
- [VSCode] Fixed a hook's reason for blocking or stopping a prompt disappearing after a window reload
- [VSCode] Fixed tabs stuck on a conversation that can't be resumed: the error now says so and offers to start a new conversation
- [VSCode] Fixed every file Read, Write and Edit stalling for ten minutes and then being skipped when the editor stops responding to the extension's automatic save before the tool runs
- [VSCode] Fixed the agent map labeling a sub-agent with the session's model instead of the model it actually ran on (e.g. under
CLAUDE_CODE_SUBAGENT_MODEL_FORCEor an agent's ownmodel) - [VSCode] Fixed uninstalling a plugin from the Manage plugins dialog, which removed the wrong installation or failed for a plugin installed for the project
- [VSCode] Fixed sign-in staying on the authorization-code step after going back and choosing the same sign-in method again
- [VSCode] Fixed the chat panel stalling when a long session trims its oldest rows
- [VSCode] Fixed the conversation disappearing from a session when many agents run
- [VSCode] Fixed the editor tab keeping an old name after a session was renamed with /rename, by a SessionStart hook, or on claude.ai
- [VSCode] Fixed the agent map's transcript view leaving out messages sent to a running agent
- [VSCode] Improved the Manage plugins dialog: a failed plugin action now opens a popup that explains it and, where there is one, offers the fix
- [VSCode] Changed the Manage plugins dialog to ask before removing a marketplace or turning off a plugin that your project's shared
.claude/settings.jsonturns on - [Cloud sessions] Fixed Run now on a routine showing internal error text when the run is refused before it starts; it now shows the same explanation as the routine's failure notification
- [Cloud sessions] Changed
MCP_DISCOVERY_CACHE=1, when set in your cloud environment's variables rather than a settings file, to reuse your connectors' tool lists after a session restart; other MCP servers are no longer cached and connect at startup - [Claude Tag] Added direct messages with Claude for members on an Enterprise plan Standard or Usage-Based Chat seat who also have Cowork; a seat that includes Claude Code is no longer required
- [Claude Tag] Fixed the Default model setting in admin settings and a channel's Configure page offering models your organization can't use, which made saves or new sessions fail
- [Claude Tag] Fixed the note under Claude's Slack messages saying it answered on a fallback model, and why, disappearing when Claude later edited that message
- [Code Review] Fixed the organization menu in Code Review's "Add a repository" dialog showing only a few of your GitHub organizations; it now loads more as you scroll
- [Code Review] Improved the Code Review check run to say when your repository's REVIEW.md wasn't applied, for example on a very large pull request or when REVIEW.md is a symbolic link
- Added
-
🔗 earendil-works/pi v0.99.1 release
New Features
- GPT-6.1 Sol — Available on OpenAI, Azure OpenAI, and OpenAI Codex, and now the default OpenAI Codex model. See Select a model.
Added
- Added GPT-6.1 Sol (
gpt-6.1-sol) to the OpenAI, Azure OpenAI Responses, and OpenAI Codex providers.
Changed
- Changed the default OpenAI Codex model to GPT-6.1 Sol (
gpt-6.1-sol).
Fixed
- Fixed
/loginwith OpenAI failing in the bundled release with a missingopenai-chatgpt.jsmodule error.
-
🔗 Evan Schwartz Please add prompt caching to Jev-style models rss
If you're building a Jev-style "System One" model, please add prompt caching or reusable question sets to your API 🙏. This would make batch use cases even more efficient, so you could amortize the cost of many questions asked over the same input. (This was also proposed in typesafe-ai/typesafe-sdk-js#10.)
TL;DR: after a week of tweaking my Jev calls, my questions are ~88% of the input tokens. I'm asking 54 questions of ~1.1 million documents per month. Jev makes certain types of classification tasks easy and cheap, but prompt caching would make batch workflows even more cost effective. For me, the total dollar amount is still reasonable (less than $150 per month), but I'm sure others will hammer these APIs even harder.
Context
I work on Scour. It’s a personalized content feed where you describe topics you’re interested in and it finds articles and blog posts related to them.
For some time, I’ve wanted to ask slightly fuzzy questions of each piece of content Scour pulls in. What level of experience does this assume? Will this still be worth reading after a week, or is it news that goes out of date quicker? However, running the >1 million pieces of content Scour ingests each month through even the cheapest LLM would be prohibitively expensive for a bootstrapped project. System One models make this kind of use case feasible.
TrAiN yOUr OwN cLaSsIfiEr
Yes. I could. I might. But having a kind of general purpose classifier that I can ask a barrage of questions and add more questions to over time without retraining is very interesting.
Tuning the question set
I’ve spent a good chunk of the last week running experiments with Jev, tuning questions, comparing Jev's answers to panels of LLM judges and spot checking the results. The questions help me weed out junk that was hard to spot deterministically, figure out the expertise assumed by posts (to hide beginner content from advanced readers), and more.
Following TypeSafe's advice about atomic questions and sending all questions in one request, v28 of my question set includes 54 questions: 50
nouls, 3choices, and 1score. I also reworded most of my questions to remove as much of the criteria description as I could without losing too much accuracy on the results.This set of questions is approximately 2,150 tokens, including Jev’s fixed overhead that seems to be around 260 tokens (judging from the token count for a trivially short question). When I send this set of questions in with the post’s title, URL, and summary or first snippet of text, my questions represent ~88% of my input tokens. The questions are sent verbatim with every request.
Note that I don't send multiple posts in the same request because of the warning that "Accuracy falls as the state grows with content unrelated to the decision."
Prompt caching or "Register-once questions"
To TypeSafe and all the other labs that follow with these types of models, please add prompt caching or the ability to register sets of questions and criteria and reuse them.
I don't know enough about how these models are served and whether there are caching efficiencies to be gained on the backend, but this would help increase Jev's effective "intelligence-per-dollar".
With cheaper repeated questions, I'd put back some of the criteria I cut down, and probably add even more questions.
As another alternative, an API that takes multiple questions and multiple inputs and gives you the results for each question evaluated for each input would give some of these savings without taking the accuracy hit from packing multiple inputs into one request.
Batch mode?
Async batch mode probably makes sense for super high-volume, fully offline use cases. It's less relevant for me because I want newly ingested content to be available soon after Scour reads it. But others will probably find plenty of use for it.
Thanks
Keep up the great work! This is a very exciting development.
-
🔗 earendil-works/pi v0.99.0 release
New Features
- Codemode and MCP — Connect MCP servers and let models run JavaScript that calls tools in parallel. See MCP Servers and Enable codemode.
- System theme — Pi's colors now come from your terminal's own palette by default. See Use your terminal's colors.
- Sign in with ChatGPT — Use a ChatGPT subscription with the OpenAI provider through
/login openai. See Authenticate interactively. - Virtual models — Extensions can route each request to a different physical model. See Virtual Models.
- Classifier models — Run Jev classifiers from codemode scripts, or use any llama.cpp model as a classifier. See How codemode works and Classification.
Added
- Added codemode, tool search, and MCP support as built-in extensions. The
codemodetool runs model-written JavaScript in a QuickJS sandbox that calls pi's tools; enable it withdefaultToolsor--toolsand configure it withcodemode.modeandcodemode.inlineBudget.tool_searchfinds tools that are not declared to the model and declares them. MCP servers over stdio or streamable HTTP, with OAuth, come frommcp.json(global, or per project once trusted) orpi.registerMcpServer()and are managed with/mcpandpi mcp add|remove|list|login|logout. See MCP Servers and Enable codemode (#10040). - Added extension tool APIs for orchestrating tools:
exposure(direct,model-only,codemode,deferred, orhidden),namespace,annotations,outputSchemawithstructuredContent,isErrorresults,prepareLoadout(), andctx.executeTool()for nested tool calls, which emit events withparentToolCallIdand are recorded as boundednestedCallson the calling tool's result. See Tool exposure. - Added a warning when an extension that registers the same tool, command, or flag replaces a built-in extension (#10174 by @cristinaponcela).
- Added experimental virtual models: extensions register them with
pi.registerVirtualModel()and pick a physical model and thinking level for each request. The footer shows the routed model,/sessionlists cost per physical model, andexamples/extensions/jev-router.tsroutes with the Jev classifier. See Virtual Models. - Added Sign in with ChatGPT to the OpenAI provider in
/login, which uses a ChatGPT subscription with the OpenAI API. Pi stores a stabledeviceIdin the global settings for this login and omits it from bug reports. - Added the
systemtheme, now the default, which derives pi's colors from the terminal's reported foreground, background, and ANSI palette and rebuilds them when the terminal switches between light and dark. See Use your terminal's colors. - Added
#rgb,oklch(), andokhsl()colors and an optionalappearancefield to theme files, andtheme.style(),theme.colors, andtheme.appearancefor extensions. See Themes and TUI. - Added a classifier model for every llama.cpp chat model, answered from next-token label probabilities. See Classification.
- Added inherited Jev classifier models on OpenRouter, Cloudflare Workers AI, Vercel AI Gateway, and OpenCode Zen.
- Added the
fullscreenWheelScrollLinessetting and/settingsentry for fullscreen mouse-wheel scrolling. The default"auto"accelerates fast wheel spins outside local macOS terminals (#9758). - Added per-input disposition to successful RPC
prompt,steer, andfollow_upresponses,AgentSession.steer()/followUp(), andRpcClient.prompt()/steer()/followUp();RpcClient.prompt()also acceptsstreamingBehavior(#9098, #9803). - Added image generation to
ModelRuntime:generateImages()with runtime-resolved auth (stored credentials, OAuth, runtime API keys,models.jsonheaders), plusgetModelsOfType(),getModelOfType(),getAvailableOfType(),getAllModels(), andgetAllAvailable(). OpenRouter image models are listed under theopenrouterprovider and share its credential; an upstream ID can have separate chat and image entries.models.jsonproviders and extension registrations without a model list keep built-in image generation. Extension model lists can include discriminated chat, image, and classifier entries with operation implementations; when supplied, they replace the provider catalog across every operation. Chat-facing reads (getModels(),getAvailableSnapshot(), the model picker) are unchanged. - Added classifier support to
ModelRuntime, includingclassify(), classifier model accessors, runtime-resolved authentication, and the built-in TypeSafejev-latestmodel. - Added
types=chat,image,classifierto pi.dev model catalog requests so remote refreshes overlay every supported model type; entries of unknown model types are ignored. - Added the
provider_stream_eventextension event for observing parsed provider events before normalization, with an opt-in/debug-providerexample viewer (#9784, #9901 by @davidbrai). - Added a show/hide toggle (
H) in HTML exports for custom messages markeddisplay: false. Messages remain hidden by default and can also be revealed from the sidebar (#8896, #10020 by @rwachtler). - Added inherited Claude Sonnet 5.5 support for Anthropic with adaptive thinking and a 1M context window.
- Added a Built-in section in
pi configto disable the built-inmcp,llama.cpp,codemode, andtool-searchextensions globally or per project, stored as-builtin:<name>in theextensionssetting. SDK inline extensions opt in withbuiltin: true. - Added
+nameand-nameentries to thedefaultToolssetting to add or remove tools without repeating the defaults, for example"defaultTools": ["+codemode"]. Project entries of this form apply on top of the user setting. Documented how to enablecodemodewithout MCP and how to use classifier models such as Jev from codemode scripts. - Added the token usage and cost of codemode
models.classify()calls to the codemode tool result, so they count toward the session cost; the codemode result shows each call's cost.
Changed
- Switched the build from the TypeScript native preview to TypeScript 7.0 with an ES2024 target, and replaced
tsxwith Node's built-in type stripping for running from source (#9965). - Removed the
[Themes]section from the startup banner. Custom themes remain available in/settings, and theme conflicts are still reported. - Changed the startup header to show the pi logo with the version instead of the app name.
- Changed the built-in
darkandlightthemes to the revised pi colors, written in OKHSL. - Changed light/dark terminal detection to use the reported background color first, then the terminal's light/dark report, then
COLORFGBG. The first-time setup no longer shows the detected appearance. - Renamed the inherited OpenAI Codex provider to "OpenAI Codex (legacy)"; Sign in with ChatGPT on the OpenAI provider supersedes it.
- Changed inherited terminal detection to treat
TERM=*-directas truecolor. - Built-in extensions and tools are named
builtin:<name>(for examplebuiltin:mcpandbuiltin:read) in errors, diagnostics, RPC source info, and bug reports, instead of<inline:name>and<builtin:name>. Their slash commands no longer carry a[t]autocomplete tag. --no-extensionsalso disables the built-in extensions, including the llama.cpp provider. Load one explicitly with-e builtin:<name>, for examplepi -ne -e builtin:mcp.- Tool calls without a custom call renderer, including direct MCP tool calls, now show their arguments: as
key=valuepairs on the title line when collapsed and onekey: valueline per argument when expanded. MCP calls are titledserver/tooland their results collapse to 5 lines. bashandpowershellstructured results, which codemode scripts receive, now hold up to 1 MiB of output instead of the model-facing 2000 lines or 50KB, and addtruncatedandfull_output_path. Longer output keeps its first and last 512 KiB. Empty output is""instead of(no output).
Fixed
- Fixed X11 clipboard text being misidentified as an image when the clipboard owner accepts unadvertised image targets (#9786).
- Prevented managed git packages from automatically installing Pi peer dependencies and added warnings for extension packages that list host-provided modules in
dependencies(#9863). - Fixed pinned git extensions loaded with
-econtinuing to use the first downloaded commit after the ref changes (#9982). - Fixed
RpcClientskipping the next event listener when a listener unsubscribes while handling an event, which could makewaitForIdle()time out aftercollectEvents()(#9990). - Fixed full-file
readcalls rendering as:1when models sendnullfor omittedoffsetandlimit(#9996). - Fixed new sessions being lost when pi exits before the first assistant response. The session file is now created when the first user message is sent (#10000).
- Fixed unloaded llama.cpp autoload presets overwriting a cached runtime context window with the GGUF training context (#10077, #10158 by @cristinaponcela).
- Fixed custom themes ignoring
terminal.trueColorand other terminal capability overrides and rendering with 256 colors (#9973, #10039 by @christianklotz). - Fixed pasting files copied in Finder inserting the file icon image instead of the file paths; paths are quoted in bash mode (#9999, #10136 by @christianklotz).
- Fixed the startup header, loaded resources, and chat notices keeping their old colors after a theme change.
- Fixed the Fireworks default model pointing at the removed Kimi K2.6 model; it now defaults to Kimi K3.
- Fixed the OpenCode Go default model pointing at the removed Kimi K2.6 model; it now defaults to Kimi K3.
- Fixed the Together default model pointing at the removed Kimi K2.6 model; it now defaults to Kimi K3.
- Reduced CPU use while streaming in long sessions and when previewing themes: the footer caches session usage totals, collapsed bash results cache their preview, and
sanitizeBinaryOutput()no longer splits output into per-character arrays. - Fixed the usage of tools called through
ctx.executeTool(), for example from codemode scripts, being dropped from the session cost; it is now added to the calling tool's result usage. - Fixed inherited
/skillautocomplete appearing empty when loaded skill names did not contain the letters inskill(#9944). - Fixed inherited path and
@autocomplete not working after opening wrappers such as(,[,{,<, or a backtick. - Fixed inherited image stretching in terminals that use the Kitty graphics protocol (#8938, #9957 by @rwachtler).
- Fixed inherited shell cursor staying hidden after exit when an extension closed an overlay during shutdown (#10026).
- Fixed inherited keyboard input being lost after a mouse click in a
/settingssubmenu closed it. - Fixed inherited 1-hour Anthropic cache writes through Vercel AI Gateway being priced at the 5-minute rate (#9210).
- Fixed inherited model-level
samplingParamsbeing dropped by directstream()/complete()calls on OpenAI-compatible APIs (#9506). - Fixed inherited Mistral GLM requests failing with "Expected at most one leading ThinkChunk" after empty content deltas (#9674).
- Fixed inherited OpenAI Fast mode requests being priced at the standard rate (#10034).
- Fixed inherited Mistral reasoning models ignoring the requested thinking level (#9678).
- Fixed inherited OpenCode Zen and OpenCode Go
qwen3.8-flashthinking being replayed as plain text on later turns (#10047). - Fixed inherited OpenAI Responses streams from servers that omit
output_index, such as llama.cpp, running mixed-up tool calls; such streams now end with an error (#9974). - Fixed inherited Anthropic and OpenAI Codex browser sign-in waiting indefinitely after the provider redirected with an authorization error, and Anthropic sign-in failing when its callback port is in use.
- Fixed inherited GitHub Copilot Claude Opus 5.5 offering unsupported thinking levels when upstream model metadata is incomplete.
-
🔗 HexRaysSA/plugin-repository commits sync repo: +1 plugin, +2 releases rss
sync repo: +1 plugin, +2 releases ## New plugins - [mcrit-ida](https://github.com/familiary/mcrit-plugin) (2.0.0) ## New releases - [ida-mcp](https://github.com/hexrayssa/ida-mcp): 20260929.0.1 -
🔗 Simon Willison OpenAI DevDay 2026 live blog rss
I'm at OpenAI DevDay today, in Fort Mason, San Francisco. Same as last year I'll be live blogging the keynote and some other notes during the day.
OpenAI gave me a free ticket and a seat in the "creator" area for the keynote.
You are only seeing the long-form articles from my blog. Subscribe to /atom/everything/ to get all of my posts, or take a look at my other subscription options.
-
🔗 smol-machines/smolvm smolvm v1.20.2 release
What's Changed
- Resize CPUs, RAM and disks of a running machine in one command by @BinSquare in #1466
- Let a stopped machine's egress policy be replaced by @BinSquare in #1449
- Give a machine that is still starting time before the ephemeral sweep reaps it by @BinSquare in #1470
- Bump the Nix package to 1.20.1 by @BinSquare in #1469
- End a foreground packed run's VM with its CLI, and watch its mounts from the CLI by @BinSquare in #1467
- Bump the workspace to 1.20.2 by @BinSquare in #1468
Full Changelog :
v1.20.1...v1.20.2 -
🔗 smol-machines/smolvm smolvm v1.20.1 release
What's Changed
- Bump the Nix package to 1.20.0 by @BinSquare in #1463
- Count file-backed cache as headroom for live RAM growth on macOS by @BinSquare in #1464
- Bump the workspace to 1.20.1 by @BinSquare in #1465
Full Changelog :
v1.20.0...v1.20.1 -
🔗 backnotprop/plannotator v0.27.22 release
Follow @plannotator on X for updates
Missed recent releases? Release | Highlights
---|---
v0.27.21 | Remote and phone sessions load several times faster, real Request changes on GitHub, model pickers show real names, OpenCode fixes
v0.27.20 | Mistral Vibe support, annotate gets the full Options menu and Settings, jj Commits panel, long lines wrap in plan code blocks
v0.27.19 | Before/After image previews in code review, file comments as GitHub file threads, forge-correct#123links,/plannotator-lastfinds the right session
v0.27.18 | Model pickers from your installed Claude and Codex (Opus 5.5, Fable 5.1, GPT-6), unsent PR review comments survive new pushes
v0.27.17 | Diagram files open in the diagram viewer, OpenCode switches model with agent, idle review stops polling the git remote, Tree is the default review view
v0.27.16 | Themed diagrams on Mermaid 12, comment on any node or edge, patch-file review, embedded HTML documents render
v0.27.15 | Plannotator TUI and Herdr Annotate announcement, element context on pinpoints, HTML links open as linked documents, All files panel, Classic diff default
v0.27.14 | Pi plan progress survives compaction, Codex threads across rollout files, WSL browser setting, Mod+E edit mode
v0.27.13 | Open a review on a specific base (--base,--diff-type), symlink containment on /api/doc, CI flake fix, Amp decision relay
v0.27.12 | Unified decision control, token hover cards, local-vs-remote diff, approval notes
v0.27.11 | OpenCode server leak fix, durable local feedback archive, unknown-subcommand fix
v0.27.10 | Auto-viewed files on scroll, annotation undo/redo, OpenCode 2 slash commands restored, npm 12 agent terminal fixWhat's New in v0.27.22 A fix release. Plans can open the documents they link to, Claude review jobs are locked down more tightly, Code Tour works on Linux machines where Claude Code's sandbox cannot start, and Pi no longer reviews an outdated plan. Eight pull requests, two of them from community contributors, answering four reports. Plans can open the documents they link to Claude Code saves plans in ~/.claude/plans/, outside your project. Plannotator only looked inside the project for linked files, so a plan that linked or [[notes]] to a file saved beside it showed a link that went nowhere. Plan review now also looks in the plan file's own folder, but only for files the plan actually links to, matched by exact name. Every plan you have ever made lives in that folder, so a plan can open its own notes without being able to reach any other plan. Links inside code blocks do not count, and plans kept inside the project still fall back to the project root as before. If a file beside the plan has the same name as a project file, the one beside the plan wins. (#1620, closing #1443, by @bendrucker) Claude review jobs are locked down
Code review, Code Tour and Guided Review run Claude Code with a list of read- only commands they are allowed to use. That list had grown looser than its purpose: a few entries accepted options that could run other programs or write files, and the review prompt contains the diff under review, which is untrusted text. For most people, Claude Code's own sandbox is off, so that list was the only thing standing between a hostile diff and the machine.
The list is now narrowed to the commands the prompts actually use, shared by all three job types, with the risky options explicitly refused. Jobs also no longer read a repository's own
.claude/settings.json, so a pull request you are reviewing cannot add permissions or hooks to the job. Your personal Claude Code settings and any managed settings still apply. Jobs no longer load your MCP servers either; before, a Code Tour could wander into tools from your IDE.Code Tour on Linux when Claude Code's sandbox cannot start
If your Claude Code settings enable its sandbox and the machine lacks
bubblewrapandsocat, the sandbox fails to start and Claude refuses every command. Code Tour then ended with "Tour blocked: git access is unavailable" and no hint why.Now, when a Claude job has every ordinary git command refused, the Agents tab shows a warning explaining what happened. You can install those two packages, or turn Claude's sandbox off for Plannotator's jobs only with
PLANNOTATOR_CLAUDE_SANDBOX=0or{ "claudeSandbox": false }in~/.plannotator/config.json. Plannotator never turns the sandbox off on its own, since that would quietly override a choice you made.(#1628, closing #1627, reported by @Infeligo)
Pi reviews the plan you just wrote
Pi runs all the tool calls in one assistant message at the same time. When the agent edited the plan and submitted it in the same message, the submit could read the file before the edit landed, so the review opened on the previous version and version history recorded it. The submit tool now tells Pi to run its batch in order, so the edit always finishes first. Every Pi version the extension supports honors this.
(#1623, closing #1622, reported by @bugs-wkettlitz)
Additional Changes
- Claude reviews no longer fail after producing findings. Recent Claude Code versions can end one run with several result messages when the model uses background subagents, and the last one is often empty. Plannotator read only the last message, so about one code review in three was marked failed even though Claude had returned findings. Review, Code Tour and Guided Review now read the newest message that carries results (#1630).
- VS Code tabs follow a changed data folder. Plannotator looked up VS Code's connection file once at startup, so setting
PLANNOTATOR_DATA_DIRlater made opening a review in VS Code quietly fall back to the browser. It now looks the file up each time (#1621, closing #1502, by @gwynnnplaine). - Tests stay out of your data folder. Running the test suite from a package directory wrote drafts, history and saved plans into the real
~/.plannotator, and a personal config could make tests fail. Every package now uses the same temporary data folder the root run does (#1624, #1626).
Install / Update
macOS / Linux:
curl -fsSL https://plannotator.ai/install.sh | bashWindows:
irm https://plannotator.ai/install.ps1 | iexClaude Code Plugin: Run
/pluginin Claude Code, find plannotator , and click "Update now".Pi: Update
@plannotator/pi-extensionto 0.27.22 and restart Pi.OpenCode: Clear cache and restart:
rm -rf ~/.bun/install/cache/@plannotatorWhat's Changed
- fix(server): resolve the VS Code IPC registry path per call by @gwynnnplaine in #1621
- fix(plan): open docs linked beside the plan file by @bendrucker in #1620
- fix(pi): run plannotator_submit_plan sequentially so it never reads a stale plan by @backnotprop in #1623
- test(pi): sandbox review-server tests from the user's config by @backnotprop in #1624
- test: sandbox the data dir when tests run from a package directory by @backnotprop in #1626
- fix(agents): hermetic Claude review jobs + sandbox opt-out for Linux by @backnotprop in #1628
- security(agents): narrow Claude job command allowlists to what the prompts use by @backnotprop in #1629
- fix(agents): read structured output from any Claude result event, not just the last by @backnotprop in #1630
Contributors
@bendrucker made plans able to open the notes and documents saved next to them, a gap every Claude Code user with multi-file plans ran into.
@gwynnnplaine tracked down the last module that froze the data folder at startup and fixed it with tests that reproduce the bug.
Community
- @bugs-wkettlitz diagnosed the Pi race down to pi's parallel tool execution and proposed the exact fix in #1622
- @Infeligo reported Code Tour failing on Linux, with the full tool log, in #1627
Full Changelog :
v0.27.21...v0.27.22 -
🔗 HexRaysSA/plugin-repository commits sync repo: -1 plugin, -4 releases rss
sync repo: -1 plugin, -4 releases ## Removed plugins - mcrit-ida -
🔗 smol-machines/smolvm smolvm v1.20.0 release
What's Changed
- Bump the Nix package to 1.19.3 by @BinSquare in #1441
- Add opt-in exact and wildcard egress host patterns by @ankrgyl in #1438
- Keep a paused machine's disks in place, and share backing copies across pauses by @LoganGrasby in #1442
- Stream sparse RAM into stored checkpoints by @BinSquare in #1443
- Quote --at '~N' in hints and docs so zsh doesn't expand it by @BinSquare in #1446
- Stop decoding pause archives after resume inputs by @BinSquare in #1445
- Remove duplicate reads and RAM copies from pause resume by @BinSquare in #1447
- Parse checkpoint indexes from memory so saves and restores stop slowing down as history grows by @BinSquare in #1448
- Layer a macOS branch over a restored disk with an absolute backing path so it can boot by @BinSquare in #1451
- Bump libkrun to live resource resizing and branch pause fixes by @BinSquare in #1450
- Create machines from an unchanged private pack without re-reading it by @BinSquare in #1455
- Stop re-extracting VM-mode packs on every launch by @BinSquare in #1456
- Let a non-root resume ignore the root-only restore tmpfs it cannot read by @BinSquare in #1452
- Bump libkrun to the macOS branch pause fix by @BinSquare in #1460
- Re-extract a pack entry whose image layers were removed after extraction by @BinSquare in #1459
- Grow running machines without rebooting by @BinSquare in #1319
- Boot the children of concurrent branches from one frozen source together by @BinSquare in #1457
- Pause packed machines in place and stop their pause from crashing the VM by @BinSquare in #1461
- Bump the workspace to 1.20.0 by @BinSquare in #1462
Full Changelog :
v1.19.3...v1.20.0 -
🔗 Filip Filmar Vreteno: a RISC-V core written in TxHDL rss
Vreteno is a 32-bit RISC-V processor core designed in TxHDL. TxHDL is a hardware description language implemented as a native Rust library. The core implements the RV32IMC instruction set architecture: the base integer set, standard multiply/divide extensions, and compressed instructions. It supports traps, hardware interrupts, and system memory access via an AXI bus. Lockstep verification checks core state against a reference model on every cycle, while gate-level netlists match cycle-accurate Rust simulations. This article details the core microarchitecture, verification strategy, and FPGA implementation.
-
🔗 Filip Filmar wdbcvt: Bundling Waveform Test Corpora and Daily Nightly Releases rss
wdbcvt now publishes two new release archives with every build: a documentation bundle and a waveform corpus. The waveform archive packages the test suite’s
.wdbfiles, ground-truth.fstconversions, and HDL source files together. Nightly builds now run daily, check the repository forge for new commits, and publish all build artifacts to Forgejo and GitHub.Why a standalone waveform corpus matters
Tool developers who study Vivado’s
.wdbformat need test cases. Vivado is a multi-gigabyte software package. Running simulations requires licenses and tool setups that many open-source developers do not maintain. Thewdbcvtrepository already contains dozens of simulated test cases across VHDL, Verilog, and SystemVerilog. Previously, these test cases stayed inside the repository tree alongside intermediate simulation files. Developers of external tools had to clone the full repository or run Vivado locally to obtain sample waveforms. -
🔗 Armin Ronacher Deser: Rethinking Rust Serialization rss
Serde is an amazing serialization library for Rust and it has been a huge reason why I felt productive with it for years. However already while at Sentry I got quite frustrated with some of the limitations with it but actually replacing Serde is tricky because of the might that it has in the ecosystem. Also because it's quite hard to actually do better without also making some potentially painful compromises.
Here are three examples of Serde corner cases that show poor interactions of Serde features or unexpected limitations:
A number that is a map
An internally tagged enum, with
serde_json'sarbitrary_precisionfeature turned on:#[derive(Deserialize)] #[serde(tag = "type")] enum Shape { Circle { radius: f64 }, } serde_json::from_str::<Shape>(r#"{"type": "Circle", "radius": 1.5}"#) // error: invalid type: map, expected f64Serde's data model has no place for arbitrary precision numbers, so
serde_jsonuses in-band signalling with a map with a magic key. The enum has to buffer the fields until it has seen the tag, and the buffer does not know about the magic key. Because Cargo features are unified, it's enough for any crate in your dependency graph to turn the feature on.Flattening breaks integer keys
#[derive(Deserialize)] struct Stats { scores: HashMap<u32, u32>, } #[derive(Deserialize)] struct Report { name: String, #[serde(flatten)] stats: Stats, } serde_json::from_str::<Report>(r#"{"name": "x", "scores": {"42": 23}}"#) // error: invalid type: string "42", expected u32 at line 1 column 35Statson its own parses{"scores": {"42": 23}}just fine. JSON keys are always strings, andserde_jsononly turns them into integers if the type asks for one. However onceflattenbuffers the value,"42"is just a string. The error also points at the end of the document rather than at the key.Adapters do not compose
fn from_hex<'de, D: Deserializer<'de>>(d: D) -> Result<u32, D::Error> { ... } #[derive(Deserialize)] struct Theme { #[serde(deserialize_with = "from_hex")] primary: u32, #[serde(deserialize_with = "from_hex")] accent: Option<u32>, } //error[E0308]: `?` operator has incompatible types // | // | #[serde(deserialize_with = "from_hex")] // | ^^^^^^^^^^ expected `Option<u32>`, found `u32` // | //help: try wrapping the expression in `Some` // | // | #[serde(deserialize_with = Some("from_hex"))] // | +++++ +A function cannot be passed as a type parameter, so there is no way to apply
from_hexto the inside of anOption, aVecor a map. You write another function for every wrapper, and once you havefrom_opt_hexthe field is no longer optional unless you also remember to add#[serde(default)].None of these are bugs that are easy to fix in Serde. They fall out of its design, and that design is protected by Serde's stability guarantees.
Back in 2022 I started an experiment called Deser. It's a serialization library for Rust that takes the user experience of Serde and puts it on top of a completely different architecture inspired by miniserde. I never really finished it and it sat around for a few years. I picked it back up, and it has now reached a point where I think it's worth looking at. Even just to inspire others to see if they want to explore the space.
The Name And Idea
The name is Serde with its two halves swapped. Deser is Serde but the other way around. In Serde, a type drives the deserialization process: a
Deserializeimpl asks the deserializer for the kind of value it expects, the format calls back into a visitor. Every nested value is handled by recursion which makes Serde deserialization inherently grow the stack with each level of nesting.Deser on the other hand turns this around and the format tells the type of the next value and pushes events into a sink. When a sink hits the start of a nested value, it doesn't call into it but hands back a new sink to a driver, which keeps all state on the heap (in fact, in an arena). On the way out, emitters return their nested values instead of recursing into them.
That also means that Deser cannot support formats like protobuf that are not self describing. They are in fact quite intentionally left out of the design entirely. Which is one way to say: if you want to "fix" Serde, you need to make some other compromises.
Most of the reasons for Deser's ideas go back to Sentry Relay, which processes enormous amounts of untrusted JSON. Over the years when I was at Sentry we ran into the same set of problems again and again, and many of them are not really bugs in Serde but consequences of its design. Serde's stability guarantees mean that a lot of them cannot be fixed without breaking every format and every hand written implementation. Most of these problems come from three decisions:
-
One set of traits for all formats. Serde serves both self describing formats (JSON, YAML, TOML, …) and formats where the reader has to know the type upfront (postcard, bincode, protobuf, …). That is incredibly useful, but it means that some features only work with some formats, and you find out at runtime. In case of Serde it also has some odd wrinkles where a derived struct quietly accepts an array in place of an object in JSON for instance.
-
A fixed data model that loses information when buffering. Internally tagged enums, untagged enums and
flattenneed to buffer values before they know what to do with them. The buffer can't hold everything the format knew, errors lose their location and extensions to the ecosystem rely on in-band signalling to express things such as arbitrary precision numbers. -
Recursion on the call stack. Every level of nesting uses stack space. Formats protect against this with a recursion limit, but the moment you go through a code path that doesn't have one (writing, dynamic values), deeply nested data can take down your process. It also means that a deserialization cannot be paused while you wait for more input.
Many of the corresponding Serde issues have been open for years, and I wrote about abusing Serde before. People have tried different angles on this over the years. Some went minimal and dropped most features to get fast compiles and no recursion. dtolnay's own miniserde is the best example of that, and deser's trait design was originally modelled after it. Other recent attempts went for runtime reflection, or for a new data model with a focus on binary formats.
If you want to read up on all of the collected challenges with Serde's design, I maintain a lengthy list here.
Dethroning Serde
First of all I don't think it's likely that one can replace Serde. The orphan rule entrenches Serde incredibly well in the ecosystem. But some things are within the reach of a crate author's control. In case of Deser it's completeness.
Deser today implements all important self describing formats from YAML, JSON, TOML, CBOR, JSON5 and the likes, but also XML and plist to really close the gap. XML in particular is something Serde has declined to support, and it shows (more on that below). At the very least format support should not be the reason not to use Deser.
The second problem usually is that actually solving Serde's issues comes at a significant cost in compile time and/or runtime performance. Deser is no different. While Deser's compile times are a bit better than Serde's, the binary bloat is quite a bit worse and the runtime performance is mixed. It's roughly comparable if you look at the numbers but depending on the format structure you are losing significantly from some of the tradeoffs.
That said, it's now in a state where it's at least in principle a drop-in replacement where the tradeoffs might work well for users.
Deser's Design
Deser does not try to be significantly different than Serde on the surface level. For most uses you derive
SerializeandDeserializeand then start using it with your format implementing crate of choice. Most attributes are very similar, though they are taking Rust expressions instead of strings.use deser::{Serialize, Deserialize}; #[derive(Debug, Serialize, Deserialize)] #[deser(rename_all = "camelCase")] pub struct Account { id: u64, account_holder: String, #[deser(default)] is_deactivated: bool, } let account: Account = deser_json::from_str(json)?;The difference in the design would become more apparent if you implement a serializer or deserializer yourself. Instead of visitors that call into each other recursively, deserializing a type creates a sink which receives events that are directly emitted by the parser, and serializing produces emitters that hand out values. Nested sinks and emitters are handed back to a driver, which keeps them on the heap. This design, which is entirely stolen from miniserde, gives some interesting consequences:
- No stack overflows. You can arbitrarily nest structures without issues. For untrusted input you set limits with a layer, and you pick the number that you are comfortable with, which is independent of your stack space.
- Suspendable. Because the state lives in the driver, a deserialization can be fed input as it arrives. It's also
Send, so it can move between threads while you wait on IO which makes it much nicer to use with tokio. Formats like JSON, CBOR and MessagePack can be parsed as a stream if you so desire. - An extensible data model. The core data model is small and made of atoms, maps and sequences. For all else, there are extension values (
DateTime,Uuid, etc.) that also all carry a fallback for formats that don't understand them. Unlike Serde this means it does not rely on in-band signalling of objects with magic keys to smuggle values through. - Lossless buffering. When a value has to be buffered (for instance because the tag of an internally tagged enum comes last), Deser records the events together with everything the format knew about them. Protocol specific extension types or error locations all survive.
- Layers are a middleware system that sit between the format and your types and can track things like paths, enforce safety limits, rename keys or redact values without having to touch specific code paths.
- Native flattening that doesn't buffer at all.
On top of that are a lot of things that I just wanted to have:
- Allow enum tags to be of any type, not just strings
- Adapters that compose (
as = Option<Vec<DisplayFromStr>>) - Enabling validation as an adapter
- derive attributes that are real Rust expressions instead of strings
- bytes as a core functionality in the data model
- duplicate keys rejected by default and errors that point at the problem
Here is a small configuration type that shows a few of these together:
use deser::adapters::DisplayFromStr; use deser::de::Recording; use deser::{Deserialize, Serialize}; use deser_encoding::Hex; use deser_validate::{Check, NonEmpty, Range}; use ipnet::IpNet; #[derive(Debug, Serialize, Deserialize)] pub struct Config { // at least one 256-bit key, each written as hex #[deser(as = Check<NonEmpty, Vec<Hex>>)] secret_keys: Vec<[u8; 32]>, // `IpNet` knows nothing about deser, but has `FromStr` and `Display` #[deser(as = Option<Vec<DisplayFromStr>>)] allowed_networks: Option<Vec<IpNet>>, listeners: Vec<Listener>, } #[derive(Debug, Serialize, Deserialize)] #[deser(tag = "type", rename_all = "snake_case")] pub enum Listener { Unix { path: PathBuf }, Tcp { host: IpAddr, #[deser(as = Check<Range<1, 65535>>)] port: u16, }, // types this version does not know are kept and written back #[deser(other)] Other(#[deser(tag)] String, Recording), }Adapters are types, so
Hexcan go inside aVec, andDisplayFromStrinside aVecinside anOption. Validators are adapters too, soCheck<NonEmpty, Vec<Hex>>decodes the keys and then checks that there is at least one. The catch-all variant keeps the tag and a recording of everything else in case someone wants to process it later.Errors are something I care a lot about, so here is what happens when a value is wrong:
secret_keys = ["9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08"] allowed_networks = ["10.0.0.0/8", "fd00::/8"] [[listeners]] type = "unix" path = "/run/app.sock" [[listeners]] host = "127.0.0.1" port = 0 type = "tcp" [[listeners]] type = "quic" host = "::1" alpn = ["h3"] let config: Config = deser_toml::Deserializer::from_str(input) .deserialize_with(|driver| driver.push_layer(PathLayer::new()))?;Note that here the tag of the internally tagged enum comes last which means that the values have to be buffered until the tag is known. In Serde this is tricky and we would lose the location if we used some tricks to add it. With Deser however, with the path layer enabled Deser you where in the structure the problem is:
Unexpected: invalid value: must be between 1 and 65535 at line 10 column 8 (path: listeners[1].port)Deser Meta Data
Deser really wants to be extensible, and XML is a more extreme example of the differences between Deser and Serde. Here is an Atom entry that mixes in Dublin Core for the authors:
use chrono::{DateTime, Utc}; use deser::Deserialize; use deser_value::Value; use deser_xml::DeserializerConfig; deser_xml::namespace!( atom = "http://www.w3.org/2005/Atom", dc = "http://purl.org/dc/elements/1.1/", ); #[derive(Debug, Deserialize)] struct Entry { #[deser(rename = atom!("title"))] title: String, #[deser(rename = dc!("creator"))] creators: Vec<String>, #[deser(rename = atom!("updated"))] updated: DateTime<Utc>, } // entries we understand, and everything else is kept as it is #[derive(Debug, Deserialize)] #[deser(untagged)] enum Item { Entry(Entry), Other(Value), } let item: Item = DeserializerConfig::new() .resolve_namespaces(true) .from_str(r#" <entry xmlns="http://www.w3.org/2005/Atom" xmlns:d="http://purl.org/dc/elements/1.1/"> <title>Deser</title> <d:creator>John</d:creator> <updated>2026-09-29T21:00:00Z</updated> <d:creator>Jane</d:creator> </entry> "#)?;XML uses namespaces which means that names need to be matched by their namespace, not by the prefix the document happens to use. Here the document says
d:and the type saysdc!.atom!("title")is just the string{http://www.w3.org/2005/Atom}title, which works because attributes are expressions. The two creators are collected into oneVeceven though there is another element between them, and the text ofupdatedgoes straight into achronodatetime. Because the enum is untagged, the entry has to be buffered before a variant is picked, and deser's buffer keeps both creators. So the result is anEntrywith John and Jane.quick-xml, the most popular XML crate for Serde, drops the prefixes and ignores namespaces entirely, so a
<x:title>from some other namespace is happily accepted as the title of the entry. The split list part though is considerably worse. A plainEntryfails with a duplicate field error forcreator, unless you turn on theoverlapped-listsfeature (which, remember, is a global additive flag that any crate could set). That feature makes quick- xml read ahead to the end of the element and buffer everything in between, without a limit unless you set one.But the feature only helps when quick-xml is hooked up to the struct directly and no buffering is taking place. Wrap the struct in the untagged enum and Serde buffers the entry itself. Read from that buffer,
Entryseescreatortwice and fails again. The fallback is a map, which keeps only the lastcreator, and there is no error. With or without the feature you get this:Other({"creator": {"$text": "Jane"}, "title": {"$text": "Deser"}, ...})Notice how John is gone.
Format specific extension types such as TOML datetimes are another case. TOML has them natively, Serde's data model does not, so the
tomlcrate passes them on as a map with a magic key. In Deser a datetime is an extension value, which formats that know it keep and all others write as a string:let value: Value = deser_toml::from_str("released = 2026-09-29T21:00:00+02:00")?; deser_json::to_string(&value)?; // {"released":"2026-09-29T21:00:00+02:00"} deser_toml::to_string(&value)?; // released = 2026-09-29T21:00:00+02:00The same with
serde_json::Valuegives you{"released":{"$__toml_private_datetime":"2026-09-29T21:00:00+02:00"}}, and reading the value into achrono::DateTimefails outright withinvalid type: map, expected an RFC 3339 formatted date and time string.The Cost
So now that you know Deser is at least in theory cool, at what cost?
It is not free. The design relies on dynamic dispatch and on sinks and emitters that live on the heap, and that has considerable runtime overhead. In my own measurements for JSON, Deser reads somewhere between 33% faster and 60% slower than
serde_json dependingon the data. On average it's about 10% slower for reading. Writes are between three times as fast and 70% slower and a wash on average. For YAML and TOML it's noticeably faster than the Serde based crates, but that is more about the format implementations than the architecture.Compile times slightly are better, but not dramatically so. Because it doesn't monomorphize everything, release builds of derived code are about 2.3 times as fast as with Serde and that get a tiny bit better in practice for your own code as less recompilation is necessary.
To make Deser's design work at all, it also uses
unsafeinternally. Most of this is to keep the chain of borrowed sinks on the heap. I feel like this is fine in the days of Miri and agents, but I know it makes some folks uneasy.And well, the biggest cost is that it's just not Serde.
How Much Is There?
Quite a lot actually which might be surprising. In addition to the core there is support for derive.
It supports all flavorts of JSON you can think of: JSON, JSONC, JSON5 and HJSON. (Fun fact here: they are all generated out of one shared parser template) For binary handling it supports CBOR and MessagePack. Additionally it does YAML 1.1 and 1.2, TOML, XML and all three flavors of Apple's plist as well as CSV/TSV, urlencoded data and environment variables. For more crazy contraptions you can attach path info or capture location data as well as support for debug printing. You can perform validation as you parse, opt into different binary encodings in addition to base64, you can bridge to serde or capture dynamic values, transcode between formats or hook it up with tokio.
For documentation see docs.rs/deser and the code itself is on GitHub alongside many examples.
-
-
🔗 Ampcode News Plaid Speed rss
Amp now supports Plaid speed for modes that use GPT-6 Astra. Plaid uses OpenAI's ultrafast speed tier, which enables inference requests to run up to 6× faster, with 6× cost per token.
Open the mode picker in the new thread dialog, choose a mode that uses GPT-6 Astra (e.g., the default High mode), and move the speed switch past Fast to Plaid.
Plaid currently works with Amp-provided inference only, as OpenAI does not yet support ultrafast for linked ChatGPT subscriptions. When Plaid is activated, subagents and non-Plaid-enabled model inference will fall back to fast or standard speed. Like fast mode, Plaid can be toggled on or off for an existing thread.
-
- September 28, 2026
-
🔗 MetaBrainz Save the date: AMA with Silona Bonewald, MetaBrainz Foundation Executive Director rss
Silona Bonewald, our new Executive Director, will be hosting an AMA (Ask Me Anything) on our forums at:
5th October 2026
17:00 Central European Summer Time (CEST)
In this forum threadPlease respect the code of conduct, as you always do!
We are aware that this time will not suit everyone's timezone, so we will open the thread a day or so before the AMA, for early questions. We will also leave the thread open for a couple of days - but don't wait too long!
You can find more information about Silona on our Executive Director announcement post.
-
🔗 Quarkslab's blog From AI Agents to RCE - Building a Vulnerability Research Workflow rss
Introduction
A security researcher spends most of the day reading code. Thousands of lines, sometimes, just to find the few that matter. Most of that reading is not where the meaningful work happens. The real work is the moment of suspicion : this length field is validated here but not there, this state can be entered twice, this loop trusts a value it should not trust. That moment takes seconds. Reaching it can take days.
An agent can work across a codebase in minutes, connecting functions scattered across files and following relationships that would take a researcher much longer to piece together by hand. What is less clear is what comes out of that volume: real vulnerabilities, or a pile of false positives to sort through.
We wanted to turn this raw power into a method. We built an agent harness for vulnerability research: the surrounding system that manages the tools, context, structure, and rules AI agents need to work on a task. In our case, that means a queryable graph over the target, a deterministic analysis core, a set of lenses that read its output, and a staged pipeline the agents work through. The harness proposes attack vectors, explores the code paths involved, and writes proofs of concept, with a researcher checking every step that matters.
The goal was not a push-button vulnerability finder, but a methodology that can be run, inspected, and repeated : the agents handle the groundwork, correlate what they find, and bring up leads, while the researcher validates them, directs the investigation, and decides what deserves more time. What we were really after was a researcher who spends the day suspecting rather than scrolling.
This article follows the workflow from the first pass over a codebase to a working exploit, showing the output of each stage along the way. We tested it on FreeRDP10, an open-source RDP client, a large target, where it uncovered two vulnerabilities that can be chained into remote code execution on the client.
The problem, and what we did about it
Late in December 2025, one person targeted nine Mexican government agencies1. Not a crew, not a state program. Over the following seven weeks they typed 1,088 instructions , which produced 5,317 commands executed across 34 separate sessions.
Claude Codehandled around 75% of the live exploitation across 305 internal servers , while a second pipeline built onGPT-4.1read the takings and tasked the next move. Close to 200 million records left the building: tax filings, the civil registry, patient data, vehicle records, the electoral roll. The forensics come from the attacker's own recovered servers.The useful number here is the ratio: roughly five commands were executed for every instruction typed. A similar pattern appeared three months earlier in a China-linked campaign2, where the model reportedly handled 80 to 90 percent of the tactical work against roughly thirty organisations.
For defenders, the important change is the amount of work one operator can sustain. The Mexico operation was run by one person , while automation covered tasks that would previously have required several skill sets. Defenders often assume that an attacker's skills set a ceiling on what they can do; agentic tooling weakens that assumption.
The campaign did not rely on novel exploitation techniques. Its significance was the scale at which one operator could carry out the work. That raises the defensive question behind this project: if automation can reduce the cost of exploiting known weaknesses at scale, can it also reduce the cost of finding and understanding new ones? We built the workflow to test that question, and pointed it at FreeRDP.
That same ability to produce work at scale creates a different problem for defenders: more output does not necessarily mean more useful findings.
curlillustrates this well34. Its confirmed-vulnerability rate fell from above 15 percent of submissions to below 5 percent, with about one report in five during 2025 amounting to machine-generated noise, each still costing hours of a volunteer team's time. The programme eventually closed. The reports were costly because many were plausible: they cited real functions and code paths and described attack scenarios that required investigation before they could be dismissed.The limiting resource is therefore researcher attention , not the number of findings a system can produce. We designed the workflow to search broadly, then spend machine time filtering and validating candidates before asking a researcher to inspect them.
What already exists, and why we built our own anyway
Several existing systems address adjacent parts of this problem, using different approaches: Google's Big Sleep for AI-assisted vulnerability research, OpenAI's Aardvark for continuous security analysis of code repositories, and the autonomous systems developed for DARPA's AI Cyber Challenge (AIxCC).
Their results show that agent-assisted vulnerability research can work. Big Sleep56 turned up a memory-corruption bug in
SQLitethat threat actors already knew about, and reported twenty more in projects such asFFmpegandImageMagick. Aardvark7, now Codex Security8, builds a threat model from a repository, tests each commit against it, and validates what it finds in a sandbox. DARPA's AI Cyber Challenge9 took a different approach, with systems designed to find and patch vulnerabilities autonomously.Each was built for a job adjacent to ours, not for ours.
Aardvark7 watches a repository as you commit to it, which is the right instinct when the code is yours and no help at all when you are auditing something you have never touched.Big Sleep5 is an internal Google tool. TheAIxCC9 systems presume a fuzzing harness already exists and are scored on running unattended; the FreeRDP surface we cared about had no harness, and we had no intention of running unattended.The deeper reason is simpler. A tool that decides everything by itself hands you a conclusion and nothing else. You then have to redo the work to find out whether it is true, which costs roughly what finding it would have cost.
curl's maintainers spent 2025 doing that with reports from strangers. We had no intention of doing it with reports from our own tooling.The workflow, in one picture
The architecture separates two properties we wanted to use differently: language models are useful for generating and exploring hypotheses, but their outputs are not reproducible (due to their non-deterministic nature).
The system therefore has two layers. A deterministic core indexes the code into a queryable graph, runs static taint from attacker-controlled sources to dangerous sinks, and passes the result through a set of narrower lenses: known
CVEandCWEpatterns, and a handful of checks specific to the target. An agent layer above it forms hypotheses, reads what the core has surfaced, judges what is real, and drives it towards a PoC.The core is tuned for recall rather than precision. It over-produces on purpose, because a lead we never generate is one we can never recover, and the agent layer exists to make that over-production affordable. It is not there to find bugs. It is there to absorb the noise a broad search inevitably produces , so that a researcher sees only what survived.
The core is reproducible. The agent layer is not, and does not need to be. It needs to be auditable.
Concretely, the pipeline runs in eight stages:
graph -> recon -> slicing -> analysis -> triage -> poc -> chain -> exploitStage | What it does
---|---
graph| Indexes the codebase.
recon| Turns structure into attack surface.
slicing| Narrows things to a single hypothesis - no agent reasons well across an entire repository.
analysis| Runs taint, pattern matching and the other lenses over that slice.
triage| Sorts the output into real , unclear and noise.
poc| Attempts to make a survivor fire.
chain| Asks whether two weak primitives amount to one strong one.
exploit| Turns a validated primitive or chain into a working exploit.The researcher works between the stages, never inside them. Which also means they can enter at any of them: run the whole thing end to end, stop after
reconand take the hypotheses elsewhere, or hand the pipeline a finding of their own and let it do thetriageand thepoc. No stage assumes the one before it was run by us rather than by a person.The rest of the article applies these stages to FreeRDP and shows the artefacts produced at each step.
Dive into the workflow through the FreeRDP example
We pointed the workflow at the FreeRDP10 tree at commit
993499447and gave it nothing else to go on. No list of past FreeRDP CVEs, no fuzzing corpus, no hint about which subsystems have historically been soft.git clone, and go.FreeRDP is the library most Linux RDP clients are built on:
Remmina,GNOME Connections,KRDC, and the RDP support in Apache Guacamole'sguacd. The same tree also implements the server side, which is what GNOME Remote Desktop and KRDP use, so parsing code shared between the two roles is exposed from both directions. We audited the client, where every byte parsed from the network ultimately comes from a server the user has chosen to trust.Step 1 - graph, recon and slicing
graph (*) -> recon (*) -> slicing (*) -> analysis -> triage -> poc -> chain -> exploit (*) Steps explained belowGraphindexes the target into a representation the agents can query: 13,506 functions, 3,105 types and 31,705 call edges in this run, including synthetic edges for resolved function pointers.The graph is not meant to be explored as a whole. Its value is in narrowing the view.
libfreerdp/core/nego.c, for example, handles the negotiation exchange at the beginning of an RDP connection. Resolving that file gives the agent its definitions, parameters and relationships to the rest of the codebase.
From there, reachability queries turn an exposed entry point into a working set:
$ sift graph reachable nego_recv --files libfreerdp/core/nego.c libfreerdp/core/tpdu.c libfreerdp/core/tpkt.c … 6 moreStarting from
nego_recv, the working set comes to nine files. Other entry points are broader:rdpgfx_recv_pdureaches 67 files andrdp_recv_pdureaches 65. Slicing uses those relationships to keep each investigation small enough for an agent to reason about without carrying the whole repository in context.Reconadds exposure information. It groups the code into components, estimates how directly attacker-controlled data reaches each one and produces a threat model, a barrier map and an ordered slice list.
The run produced 304 slices. The two findings discussed below came from slice 166,
libfreerdp/core/nego, and slice 109,channels/urbdrc/client/data_transfer. A purely mechanical ranking would not have selected either one early. The researcher used the ranking as a map, not as a verdict.Step 2 - analysis
graph -> recon -> slicing -> analysis (*) -> triage -> poc -> chain -> exploitEach slice is analysed independently. The deterministic core builds a deeper graph, runs static taint and applies the configured lenses. Taint runs twice: first with sources and sinks already known to the engine, then again after an agent has identified target-specific vocabulary such as FreeRDP stream operations, read macros and allocation wrappers.
In slice 166 the engine produced 14 findings. Two became interesting together:
f-007 attacker-controlled length stored in nego->RoutingTokenLength &varr ? f-005 same field later used as a Stream_Write sizeThe engine could establish both endpoints but not the relationship across the lifetime of the
negoobject. It reported them separately instead of inventing a data-flow edge. This is the kind of ambiguity the next stage is meant to resolve.
Slice 109 produced 116 findings. Two pointed to a USB-redirection completion path where an error leaves an attacker-supplied
OutputBufferSizein place even though no data has been written. The client later uses that size when emitting the response, exposing bytes left in the heap allocation.At this point, the two slices have produced 130 findings, but these are still candidates, not confirmed vulnerabilities. False positives are expected at this stage, and each finding still needs to be investigated and validated.
Step 3 - triage and PoC
graph -> recon -> slicing -> analysis -> triage (*) -> poc (*) -> chain -> exploitTriage groups related findings before classification. In this run, 130 findings became 36 clusters: 20 real and reachable, 3 latent, 3 mixed and 10 false positives.
The classifications are then checked dynamically. For each candidate, the pipeline builds a small harness against FreeRDP and runs it under AddressSanitizer. This is where a plausible static finding either becomes measurable or gets discarded.
For slice 166, triage brings
f-007andf-005back together.f-007ends atnego_set_routing_token, where the length derived from network data is stored innego->RoutingTokenLength.f-005picks up the same field later, when it is passed as the length argument toStream_Writeinnego_send_negotiation_request.Connecting the two findings establishes a candidate path, but not yet a vulnerability. There are two ways for
nego->RoutingTokenLengthto be populated. The first turned out to be safe: a protocol length field limits the routing token to a size that cannot overflow the destination buffer. The PoC confirmed the bound, so that path was discarded.The second path comes from a Server Redirection PDU. A malicious server can provide
LoadBalanceInfo, which is later reused as the routing token when FreeRDP establishes the redirected connection. This path does not have the same length restriction.We reproduced it using a malicious server and an ASan-instrumented FreeRDP client. The server sent a 600-byte
LoadBalanceInfopayload, which became a 615-byte routing token once wrapped with the cookie header. AtStream_Write, FreeRDP attempted to copy those 615 bytes into the 512-byte allocation created bynego_send_negotiation_request:==4514==ERROR: AddressSanitizer: heap-buffer-overflow WRITE of size 615 #2 Stream_Write #4 nego_send_negotiation_request libfreerdp/core/nego.c:1098 #9 rdp_client_redirect libfreerdp/core/connection.c:715 0x75e976227180 is located 0 bytes after 512-byte region [0x75e976226f80,0x75e976227180) allocated by thread T2 here: #1 Stream_New #2 nego_send_negotiation_request libfreerdp/core/nego.c:1083The
rdp_client_redirectframe confirms that the overflow is reached through the server-driven redirection path rather than by calling the vulnerable function directly from the harness.The test used
WITH_VERBOSE_WINPR_ASSERT=OFF, the configuration used by Debian and Ubuntu builds. With verbose assertions enabled, the bounds check aborts the client before the out-of-bounds write reachesmemcpy.The USB-redirection finding went through the same validation process. In that case, the candidate primitive was an uninitialised heap disclosure. We configured ASan to fill new heap allocations with
0xcd; the same pattern appeared unchanged in the PDU emitted by FreeRDP:[poc] heap pattern bytes (0xcd) in region [36..100): 64 / 64Only findings that survived this dynamic validation were passed to the chaining stage.
Step 4 - chaining
graph -> recon -> slicing -> analysis -> triage -> poc -> chain (*) -> exploitThe chaining stage works only on validated clusters. Here, the negotiation bug provides an attacker-controlled but blind heap write, while the USB- redirection bug provides visibility into heap contents. The leak operates after the session is established, and a server redirection can then force the client to reconnect and trigger the write. Together, these primitives made remote code execution plausible, but not yet demonstrated.
Step 5 - exploitation
graph -> recon -> slicing -> analysis -> triage -> poc -> chain -> exploit (*)Exploitation changed the balance between the agent and the researcher. Earlier stages produce artefacts that are cheap to check: a path, a source location, an ASan trace or a measurement. Exploitation is less forgiving. A failed heap- layout hypothesis, for example, often says little about why it failed.
So we gave the exploitation stage its own harness rather than letting a general-purpose agent iterate freely. It provides exploitation-specific knowledge and breaks the process into a set of gates that must be validated before the agent can move on.
This kept the agent useful on bounded tasks: interpreting leaked bytes, enumerating objects reachable from the write, testing heap-layout assumptions and generating harnesses for individual experiments. It also made its progress inspectable. A successful step meant that an expected property had been measured, not simply that the model considered the approach plausible.
There were still limits. During this run, the agent did not converge on the complete exploit by itself. We intervened twice when the models we used stopped making useful progress: first while deriving the libc base from the leak, and later while composing the primitives into a working payload.
The exploitation proceeded through the following stages:
✔ Lab set up: client VM, hostile server ✔ Cluster 2 write primitive measured ✔ Control-flow hijack demonstrated ✘ libc base from the cluster 28 leak &boxh&boxh researcher intervention 1 &boxh&boxh ✔ libc base recovered ✔ ASLR defeated ✘ Working payload &boxh&boxh researcher intervention 2 &boxh&boxh ✔ Leak and write composed into one session -> /tmp/pwned
Conclusion
On FreeRDP, the workflow produced 130 findings across the two slices discussed in this article. Triage reduced them to 36 clusters, 20 of which were reproduced as real and reachable. Two findings provided the primitives used for client-side remote code execution.
The workflow did not eliminate noise, nor was it designed to. The deterministic core could search broadly because the researcher was not expected to inspect every finding it produced. Slicing kept individual investigations bounded, clustering brought related evidence together, and PoCs provided a way to reject or confirm the resulting hypotheses before more time was spent on them.
The slice ranking also showed why we kept a researcher in the loop. The two relevant slices were ranked 166th and 109th out of 304. The ranking could account for properties such as parsing volume and exposure, but not every reason a component might be interesting to a security researcher. It provided an ordered search space rather than deciding what should be investigated.
From
git cloneto the report sent to the maintainers took four days on a roughly half-million-line C codebase that has been audited and fuzzed for years. The model was only one part of that process. Around it were the pieces that made its output usable for vulnerability research: a queryable representation of the code, deterministic analysis, bounded slices, evidence attached to findings, dynamic validation, an exploitation harness, and researcher checkpoints.Both vulnerabilities were reported to the FreeRDP maintainers. FreeRDP was our test case; the objective was the workflow itself, and whether it could make a large codebase practical to investigate without turning the researcher's time into the bottleneck.
Advisories :
References
-
The AI-Assisted Breach of Mexico's Government Infrastructure (Eyal Sela, Gambit Security, technical report, April 2026) &larrhk
-
Disrupting the first reported AI-orchestrated cyber espionage campaign (Anthropic, November 2025) &larrhk
-
The end of the curl bug-bounty (Daniel Stenberg, January 2026) &larrhk
-
Death by a thousand slops (Daniel Stenberg, July 2025) &larrhk
-
From Naptime to Big Sleep (Google Project Zero, November 2024) &larrhk&larrhk
-
Google's AI security announcements (CVE-2025-6965, the SQLite flaw known to threat actors) &larrhk
-
Introducing Aardvark: OpenAI's agentic security researcher (OpenAI, October 2025) &larrhk&larrhk
-
Codex Security: now in research preview (OpenAI, March 2026) &larrhk
-
AI Cyber Challenge marks pivotal inflection point for cyber defense (DARPA, August 2025) &larrhk&larrhk
-
FreeRDP sources (GitHub) &larrhk&larrhk
-
-
🔗 HexRaysSA/plugin-repository commits update danielplohmann/mcrit-plugin to familiary/mcrit-plugin (reposit… rss
update danielplohmann/mcrit-plugin to familiary/mcrit-plugin (repository transferred) -
🔗 r/Harrogate Polish teachers? rss
Hi. I am looking to learn polish and I am trying to find a tutor. Can a yone get me in contact with someone who could teach me please ?
submitted by /u/honorary_guardsman
[link] [comments] -
🔗 exe.dev Pools: Our Newest Primitive rss
At exe, one of the things we pride ourselves on is developing strong primitives instead of throwing features at the wall. We still believe that the VM is a key primitive in our system. Everything is just a computer and you get to choose how you use it.
We’ve encouraged you all to build VMs, scale them, and run your production stacks on our platform. But we noticed we were missing a primitive for logically grouping VMs. Today, we’re launching pools.
A pool is a collection of VMs that share a set of resources. Before pools, a VM’s resources came directly from your plan. You could scale VMs independently, but you couldn’t define a production or staging stack with its own resources in its own region.
Pools also give you predictable pricing. We don’t want to send you a surprise invoice because one VM had a burst of traffic.
Here’s how pools work:
- Create a pool in a region, with a specific CPU and memory size.
- Move existing VMs into the pool, or create new VMs in it.
- Those VMs now share the pool’s resources.
- Set up more pools as you need them.
Pricing
Pools come with some plan changes. New customers can start using pools today, and existing customers can migrate to the new plans to get access. There are two new plans to choose from: Personal and Work.
Personal
The Personal plan starts at $15 a month for a 2 vCPU / 4 GB pool that holds up to 50 VMs. You can scale your pool up to 16 vCPUs / 32 GB of memory for $155 a month. Additional disk and bandwidth charges still apply.
We’re also introducing standalone VMs, billed at an hourly rate. Other platforms call these sandbox VMs. Customers have been asking for this for a while, and we wanted to make sure we built the right thing. You manage a standalone VM like any other VM, but you own its lifecycle and decide how to use it. You can think of it as a VM in a pool of 1.
Work
Our previous Teams plan charged per seat. This made sense at first, but customers told us they’d rather pay for compute. So the new Work plan no longer charges for seats.
The Work plan is built around pools. It has a minimum spend of $150 a month, and the first 4 vCPUs / 8 GB of memory are included. On top of that, you can create as many pools as you want, of any size, billed at a predictable rate starting at $0.105/hr per 2 vCPUs. Like the Personal plan, it also includes standalone VMs.
Companies come in all shapes and sizes. If the Work plan doesn’t fit your team, reach out and we’ll work with you to find a better solution.
Pricing changes can be hard, but we believe pools are a strong foundation moving forward. Existing customers can take their time moving over (a year, to be exact). We’re here to help, and we’re excited to see what you build.
-
🔗 anthropics/claude-code v2.1.284 release
What's changed
- Added Claude Sonnet 5.5 (
claude-sonnet-5-5), now the default Sonnet model on the Anthropic API — 1M context, $2/$10 per Mtok with $0.20/Mtok cache reads - Added a "Yes, but ask again next time" answer to auto mode's prompt before a read outside the working directories, so you can allow that one read and still be asked about later ones
- Added dollar amounts to the Claude apps gateway spend limit in
/usageand the status line (for example "$271.40 / $500.00 spent this month") when the gateway runs this version or later; the status line'srate_limits.spend_limitalso gainsused_usd,limit_usdandperiod - Added
effortSlider:decreaseEffort,increaseEffortandtoggleUltracodekeybinding actions, so the/effortslider's arrow and Tab keys can be rebound inkeybindings.json - Added
/rate-limit-optionsto/helpand the command menu for claude.ai subscribers, so the usage-limit notices that mention it point to a command you can find - Added
/mcp reconnect allin the interactive terminal to retry every MCP server that failed to connect or needs authentication at once - Added Claude apps gateway startup warnings when a managed policy's
availableModelsis empty, or leaves out the model Claude Code starts on without settingmodelorenforceAvailableModels - Added
auth: { google: {} }for Claude apps gatewaytelemetry.forward_todestinations, so telemetry can be exported straight to Google Cloud's OTLP endpoint using the gateway's Google Cloud credentials - Added certificate client authentication (
private_key_jwt) between the Claude apps gateway and its identity provider, for identity providers that issue certificate credentials instead of client secrets - Fixed a damaged response stream showing raw errors such as "JSON Parse error" or "undefined is not an object", or writing the word "undefined" into an answer, instead of being retried or reported as an interrupted response
- Fixed an overloaded or server error arriving right after a thinking block ending the turn with an error instead of being retried
- Fixed "Prompt is too long" errors that persisted after compacting: when the compacted request is still too long, Claude Code now compacts once more, keeping less of the recent conversation
- Fixed a session whose model is unavailable, with no fallback model left, showing a bare "is currently unavailable" message (or "Something went wrong" in cloud sessions) instead of the model-unavailable notice and its Learn more link
- Fixed Agent SDK sessions crashing when a user message contains an image with a malformed
source, and failing on every later turn after a malformed document block; a malformed image is now replaced with an explanatory note - Fixed MCP tool calls in a resumed session failing with "No such tool available" while their server was still connecting; the call now waits up to 10 seconds for the server
- Fixed repeated calls to the plan-usage endpoint after it rate-limits or rejects your login:
/usage,/extra-usageand IDE usage views now back off instead of re-asking - Fixed
claude mcp addreporting success when managed settings restrict MCP servers to plugins; it now refuses and says what to do, instead of saving a server that never loads - Fixed the
/pluginconfigure screen: boolean options are now a true/false choice instead of free text, number options refuse invalid input, and ←/→ change an options field instead of switching tabs - Fixed
ANTHROPIC_FOUNDRY_RESOURCEbeing interpolated into the Foundry endpoint host unvalidated; a value that is not a plain resource name is now refused - Fixed Claude Desktop behind a Claude apps gateway offering no 1M context option: the gateway now marks each 1M-capable model for Desktop automatically
- Fixed ↓ in shell mode selecting a hidden background-tasks pill, which stopped Backspace and Ctrl+U from editing the prompt
- Fixed Bash tool failing on Windows with many plugins enabled: plugin
bin/directories that don't exist are no longer added to PATH, and inherited entries aren't added twice - Fixed
sparsePathsplugin marketplaces cloning empty and replacing a working local copy on older git (before 2.39), which failed every refresh with "marketplace.json file is no longer present" - Fixed fullscreen rendering erasing the terminal output above the session when
[in transcript mode writes the conversation to scrollback (macOS and Linux) - Fixed fullscreen scroll position jumping to the previous message or to the bottom when a reply finished streaming while scrolled up
- Fixed tab bars in dialogs such as
/configand/pluginbreaking the title and tab labels mid-word in a narrow terminal; a tab that doesn't fit now moves to the next line whole - Fixed the
/modelpicker showing "+1 model" below the list after scrolling to the last model; the count now covers only the models below the visible rows - Fixed
/keybindingswriting Backspace and Delete bindings for a footer action that does nothing into the generatedkeybindings.json - Fixed a rebound agent panel close key (
footer:close) typing "x" instead of itself on the row of the agent you're viewing - Fixed vim mode
.not repeating text typed very fast (for example over ssh or in tmux) or pasted without bracketed paste, and leaving the prompt in INSERT mode after repeating a change with nothing typed (such ascwthen Esc) - Fixed vim mode leaving the cursor on an image placeholder's opening bracket after
ddon the last line oryyat the end of the prompt, whererorxwould break or delete the image - Fixed a key pressed the instant the terminal regained focus answering the Remote Control enable prompt before its short safety delay restarted
- Fixed the workspace trust dialog appearing a second time after switching renderers or updating when Claude Code was started in the home directory
- Fixed rules symlinked into
.claude/rulesfrom outside the project being skipped without ever showing the external-imports approval prompt; a.claudedirectory symlinked from outside the project now asks for the same approval - Fixed plugins from marketplaces, claude.ai and npm pre-approving their own tools via
allowed-toolsunder managedallowManagedPermissionRulesOnly; only plugins from an official Anthropic source or a source that managed settings vouch for keep that pre-approval - Fixed a failed first
claude plugin installleaving the plugin enabled and recorded when a dependency's version range could not be met - Fixed the debug log dropping a failed hook's stderr when the hook also wrote to stdout, and logging nothing for a failed hook with no output; failed hooks now also log their status code
- Fixed
{"decision":"block"}returned by Elicitation and ElicitationResult hooks being ignored; it now declines the MCP elicitation, as exit code 2 does - Fixed sessions launched without the
SendMessagetool (such as by Claude Desktop) still being told to message other sessions with it - Fixed a photo sent from the Claude app over Remote Control being lost when its queued message was pulled back into the terminal prompt to edit, and the cursor moving one character for a photo with no caption
- Fixed typing a message during an automatic usage-limit wait taking the wait out of the "Continue automatically at usage limit" setting's control when that turn hit the limit again
- Fixed usage-limit warnings suggesting
/upgradeto users already on the highest Max plan; the warnings and/upgradeitself now point at/usage-creditswhen it is available - Fixed the Explore subagent switching to Opus on the Claude API when the session runs a model ID Claude Code doesn't recognize, such as a custom model behind a proxy; Explore now inherits that model
- Fixed
/loopstatus updates in self-paced mode often not being shown because Claude wrote them only in its reasoning; Claude now writes each update, and the outcome when the loop stops, as visible text - Fixed
/ultrareviewfailing to upload the working tree when started from a git worktree that the Claude desktop app created on macOS or Linux - Fixed sandboxed Bash commands failing to start on Linux when the working directory is write-denied and contains a read-denied directory
- Fixed artifact database write results telling Claude that every viewer sees a write to a viewer's private
data/users/subtree, and added a "view" level toas_level - Fixed the Claude apps gateway answering
431 Request Header Fields Too Largeto every request from a sign-in whose identity provider lists many groups; it now accepts request headers up to 256 KiB - Improved the usage-limit wait: the limit's state and the countdown with the usage-credits option now show as one block under the prompt, and limit messages no longer repeat the countdown
- Improved the "No such tool available" error for Claude in Chrome tools called without their prefix: it now names the tool to call
- Improved Monitor event rows to show what each event printed instead of repeating the description, and stopped repeating an unchanged "Waiting for N … to finish" line after every event
- Improved Workflow tool sandbox hardening for errors thrown by async script hooks
- Improved startup time and memory use by building only the parts of the settings schema that your settings files actually use
- Improved
/claude-api:hillclimbno longer spends rounds on prompt rewordings too small for the eval to measure, and an extra page you ask for besidereport.htmlis built as one local file that loads nothing from the network - Improved lists such as
/tasks,/copyand/hooks: the details after each name now line up in one column when they fit, and otherwise sit at the right edge - Improved
claude plugin marketplace addto say when it replaces a marketplace already added under the same name from a different source, and how to undo it - Improved the startup refusal when managed settings require a sign-in (
forceLoginMethodorforceLoginOrgUUID) and an API key, token orapiKeyHelperis configured: it now names the credential in use, where it is set, and how to remove it - Improved auto-memory loading: invisible characters and tags that imitate Claude Code's own markup are neutralized in
MEMORY.mdand recalled memory notes before they reach Claude - Improved
claude remote-control: in a folder you haven't trusted yet, it now asks for workspace trust on the terminal instead of exiting - Improved artifact pages: Claude writes its design plan into the page instead of the reply, and uses the name you already gave something as the page title
- Improved the Artifact tool so that when Claude is given a claude.ai chat or project link, an artifact from a chat, or an artifact id on its own, it asks for the right link or the content instead of stopping
- Changed interactive terminal and VS Code sessions to start in auto mode when no permission mode is configured, on every plan and provider;
permissions.defaultModestill overrides it - Changed Ultracode into its own toggle in
/effort(Tab, or/effort ultracode [on|off]): it no longer forces xhigh effort and stays on at any effort level - Changed retries after a dropped connection mid-response to share one budget with the rest of the request's retries, so a failing request gives up sooner
- Changed the notice shown when a Sonnet model's safeguards flag a message to explain why it happened and to offer editing and retrying
- Changed safety-related model switches in sessions that pin an Opus model with
ANTHROPIC_DEFAULT_OPUS_MODELormodelOverrides: on the Anthropic API, the API now picks the model to switch to for each kind of flag, not the pinned model - Changed the non-interactive first turn to still wait up to 2s for connecting MCP servers named by
--allowedToolsor anmcp_toolhook, even whenCLAUDE_CODE_MCP_STARTUP_WAIT_MSis0 - Changed
/recapto decline with a short notice when it arrives relayed from a chat thread (your own included) or from a routine or webhook; typed in the terminal, the Claude apps, Remote Control,-por an SDK host, it runs as before - Changed
/artifactsto show its filter tabs beside the title with one-word labels (All, Mine, Shared), using the same tab bar as/configand/plugin - Changed artifact publishing to refuse a file on a network share (a
\\host\sharepath or a/netautomount) unless it is on a mapped network drive added with--add-dir - [VSCode] Added an optional time above each prompt and response, with a date line where the day changes (Claude Code: Show Message Timestamps setting, off by default)
- [VSCode] Added plugin load errors and notes to the Manage plugins rows, with a popup to disable, uninstall or copy the error
- [VSCode] Added an Ultracode on/off switch under the Effort slider, replacing the slider's Ultracode stop; the model pill shows "· Ultracode" at any effort level
- [VSCode] Fixed Reload Claude from the Memory dialog restarting before an edited file was saved
- [VSCode] Fixed a restored tab opening a conversation another Claude process still has open; it now asks first
- [VSCode] Fixed Focus view sections you expanded closing on their own while a sub-agent is working or when the section's first step is trimmed from view
- [VSCode] Fixed typing
/modeland Enter printing usage text into the chat instead of opening the model selector - [VSCode] Fixed
/feedbackon Vertex, Bedrock and Foundry being refused after you pressed Send; the report is now saved on this computer, as the terminal does - [VSCode] Fixed sign-in waiting up to a minute for the Python extension after a window reload
- [VSCode] Fixed Claude Code tabs that stopped responding after Restart Extensions: they now reopen on their conversation
- [VSCode] Fixed a message from another agent with no recorded sender showing as raw XML in the chat
- [VSCode] Fixed messages from other agents, sessions or channels disappearing after a reload
- [VSCode] Fixed a user's own
/mcp,/configor/settingscommand being shadowed by the extension's dialog - [VSCode] Fixed Escape stopping every background agent when no turn was running
- [VSCode] Fixed plugin install links replacing a marketplace you already have that uses the same name
- [VSCode] Fixed "Prompt is too long" errors after compaction when a large text file is attached to a message
- [VSCode] Fixed chat links to files with non-ASCII characters, spaces or brackets in their path not opening
- [VSCode] Changed
CLAUDE_CONFIG_DIRin theclaudeCode.environmentVariablessetting to apply only when it is an absolute path, and passed it to terminals that continue the chat - [Cloud sessions] Fixed a routine's Edit and Duplicate controls saying the routine was still loading while you were offline; they now tell you you're offline
- [Claude Tag] Added model family choices such as "Opus (latest)" for a thread, a channel default or your DM, so the choice follows the newest model in that family
- [Claude Tag] Added the spend that counts toward your organization-wide limit to the analytics spend projection chart, with how much of the limit is used
- [Claude Tag] Fixed the earlier Claude in Slack app's progress card and link previews omitting the repository and Create PR button when a GitHub Enterprise host name contains an underscore
- [Claude Tag] Fixed Claude staying silent in a channel whose environment declines to start it; it now posts one notice asking you to contact an admin, and retries when @-mentioned
- [Claude Tag] Changed Claude to post its private sign-in notice at every @mention from someone who hasn't connected their Claude account, instead of going quiet after the first
- [Claude Tag] Improved "Notify members now" in admin settings: one press reaches every workspace your organization claimed in an Enterprise Grid, and more members in large workspaces
- [Claude Tag] Improved Claude's wait notice on self-hosted environments with on-demand runners: it now says whether a runner is starting, a start will be retried, or no runner will start
- [Claude Tag] Improved the error shown when adding a channel manager fails because the channel's Slack workspace can't be confirmed as connected to your organization
- [Claude Tag] Improved a channel's access lists in admin settings to show the connectors, repositories and plugins an auto-join pattern attaches, and where each comes from
- [Claude Tag] Improved adding repositories as a channel manager: when your GitHub sign-in can't confirm you're a repository admin, the page asks you to sign in with GitHub
- [Code Review] Fixed Code Review giving up without posting a finished review when an unsubmitted review under its GitHub App was open on the pull request; it now retries the post first
- Added Claude Sonnet 5.5 (
-
🔗 MetaBrainz GSoC 2026: GraphQL Server For Musicbrainz rss
Hi everyone, I'm Sreehari, also known online as owlpharoah (op3kay on Matrix). I'm a second year student at IIIT Jabalpur. This summer I worked on the foundations of a GraphQL server in Rust that sits over the MusicBrainz PostgreSQL database, under the guidance of @bitmap and @jadedblueeyes.
The Setting
MusicBrainz already has an XML/JSON API, but getting related data out of it means chaining together inc parameters, and browsing support differs from one entity type to the next. Looking up five artists at once isn't really possible either.
GraphQL fixes most of this by letting a client ask for exactly the fields and relationships it wants in one query. It also lets the server check how expensive a query is before running it, instead of finding out after the database has already taken the hit.
Before writing the proposal, I built a rough prototype covering Artist, Release Group, Release, and Recording, mainly to see how things would click together. Two problems showed up: N+1 queries on relationship fields, and the fact that depth limiting alone doesn't catch a shallow query that's still expensive. Both became core parts of the proposal.
The Plan

The proposal scoped the project to six entity types: Artist, Release Group, Release, Recording, Label, and Area. Each would be queryable by MBID, with the usual relationships between them, plus aliases, tags, genres, and ratings across the board.
A few decisions were made initially:
- DataLoaders would follow a two tier split. One loader maps an entity's internal id to the hydrated entity itself and gets reused everywhere that entity shows up. A separate, thinner loader maps a parent id to a list of child ids.
- Fields that need a loader call would live behind ComplexObject, so they only run when a client actually asks for them.
- Pagination would use keyset pagination instead of offset pagination, since offset pagination gets slow and inconsistent on large tables that change often.
- Query safety would come from depth limiting plus a complexity weight on every resolver.
key goals
- A working GraphQL server covering the six entity types
- Schema level depth limiting and query cost analysis
- A performance baseline from load testing.
The Result
DataLoader infrastructure is in place across all six entities, following the two tier split. Loaders exist for tags, ratings, artist credit, genres, annotations, aliases, ISNI and IPI identifiers, and MBID to internal id resolution. Hydration loaders are shared across every relationship that points at a given entity instead of duplicated per relationship.

MBID redirect handling lives inside each loader's load function. When a primary table lookup misses, the unresolved MBIDs get batch queried against the matching *_gid_redirect table, and any hits get merged into the result map. A resolver further up never has to know a redirect happened.
Keyset pagination runs across the paginated fields using ROW_NUMBER() OVER (PARTITION BY parent_id ORDER BY child_id) in a single batched query, so one query can apply a per parent limit across a whole batch of parents. The cursor ended up as a plain integer rather than the opaque string from the proposal.

Query complexity weights reflect actual database work: a scalar field costs its default, a single hop DataLoader field costs a flat amount, a paginated one to many field scales with the requested page size, and multi hop fields carry a multiplier on top.
Integration tests cover all six entities.
Week 7's load testing with k6 turned up two findings worth fixing. There was an N+1 on the isrc field, fixed with a dedicated RecordingIsrcLoader, and a gap in the complexity limiter where a pathological query executed instead of getting rejected outright.
On the infrastructure side, CI runs pre commit hooks and no longer has dead code warnings, and documentation is wired up through Magidoc in Docker compose.
What I Learned
The two tier loader split sounded simple on paper, but it's saved me a lot of time in practice. Adding a new relationship is now mostly copying hydration logic that already works, instead of writing it fresh.
Fixture data needs to be checked against the database it's running against, not assumed to be stable. musicbrainz-docker's sample dumps import non- deterministic subsets of the data, so an MBID that resolves cleanly on my machine can point at something else, or nothing, on someone else's. That cost me a few confused debugging sessions before I figured out what was going on.
What's Next
- Criterion benchmarking, to compare the current per row queries on ComplexObject fields like Release.date against a batched loader variant, and to compare the two hop ArtistCredit and Tags loader patterns against a single joined loader.
- Moka caching is still an open evaluation.
- Extended entity coverage beyond the original six.
Conclusion
This was an awesome summer and i enjoyed thinking about and wiring up the API schema and the Postgresql database, fixing bugs, and everything in between. It was really satisfying watching the two tier loader pattern click into place once and then just work for every relationship added after it.
Thanks to my mentors for the guidance, especially on the DataLoader architecture, which I wouldn't have landed on alone. And thanks to the wider MetaBrainz community for the space to build this in. It's been a pretty nice summer of query plans, a lot of Rust, and debugging, and I'd do it again.
-
🔗 Szymon Kaliski Q3 2026 rss
Hi!
We had a great summer. Having both indoor and outdoor pools within walking distance from our home meant a lot of swimming, mostly in the kids' pool with our (now two year old!) daughter.
It's been another year where we spend the hottest months at home. The summer is pretty nice here, and most days already feel almost like vacation. We travel during the grey and cold months instead, since we don't have to worry about the school year yet.
My time was split purely between work and family, so there's not much to report, other than publishing Play with Putty at Google Labs — a research prototype exploring collaborative vibe-coding:
There's a lot more about this that's not covered by the video and I hope to share more soon. For now, you can sign up for the waitlist.
Worth Checking Out
What I've been reading lately:
- Something Incredibly Wonderful Happens — history of Frank Oppenheimer and his Exploratorium, great read, highly recommended
- Experiences In Visual Thinking — awesome walk through what it means to think visually, and exercises to get better at it
- Thirty Years that Shook Physics — a historical view of development of quantum theory
- The Character of Physical Law — a couple of very dense lectures from Richard Feynman about generic principles of (and about) physical laws
- Thinking About Mathematics — only just started this one but seems great so far, overview of philosophy of mathematics
On the web:
- Aphex Twin logo generator
- some cool prototypes around using LLMs for design explorations
- Andrew Blinn has been on a roll: 1, 2, 3, and a short talk
- Zach Lieberman on 10 years of daily sketching
- some great posts from Jamie about AI, curious to see what he'll get up to at Anthropic
-
🔗 Ampcode News Opus 5.5 rss
Claude Opus 5.5 now powers Amp's
mediummode by default, replacing GPT-5.6 Sol. If you use a ChatGPT subscription with the ChatGPT Only preset,mediumstays pinned to GPT-5.6 Sol, billed to your subscription. To switch to Opus 5.5, changemediumin Tune Modes.In our internal evals, Opus 5.5 solved 65% of tasks, up from GPT-5.6 Sol's 61% and Opus 5's 56%. And it did that for 10% less than GPT-5.6 Sol and 25% less than Opus 5.
It's still behind Fable 5.1, which solved 71%, but it costs 40% less.
More Steerable
Opus 5.5 responds much better to messages you send while it works. Add a constraint you forgot, point out a wrong turn, or correct its approach, and it folds that into the work instead of arguing or carrying on with its old plan. It doesn't drop what it was doing, either.
So don't stop it to edit your prompt and start over. Send the correction. Enter steers by default, so your message reaches it right away.
High Effort, Not Max
Opus 5.5 runs at high reasoning effort. Past high, it starts to overthink. At xhigh and max it burns through tokens, scores lower than at high, and costs several times as much. So we stop at high.
Why Not GPT-6 Sol?
OpenAI shipped GPT-6 Sol an hour after Opus 5.5. We ran both.
In our evals, GPT-6 Sol scores about the same as GPT-5.6 Sol, at half the cost. In real work, though, it's much more jagged: strong on one task, off on the next.
mediumcarries most threads, so it needs a model you can count on across all of them.When To Turn the Dial Up
mediumis the right place for most work now. Two kinds of task still belong higher:Lots of unknowns. Use
ultrawith Fable 5.1 when nobody knows the path yet. Alex had to move more than 5,000 users of an enterprise customer's SSO to new email addresses. He started the thread with:We've never migrated… from one set of email addresses to the next before… and we don't really have a way to test it.
It mapped how sign-in works today, planned the switch, and watched production afterward. The first 41 sign-ins after the switch all worked.
Taste and judgment. Use
highwith GPT-6 Astra orultrawith Fable 5.1 when the hard part is deciding what good looks like. Ask for options, not one answer. Hamish wanted a better transition out of the iOS image viewer:Right now it's a hard cut, so can you please give me some options for animations and record them and show me them in line in the transcript here.
It recorded four options on the simulator. He replied "Ship B to main."
Turn the dial up when a miss costs more than the wait, not because the task is long.
How To Use It
Opus 5.5 is persistent. It keeps going until the job is done, which means it needs to know what done is, and it needs a way to see for itself. These are the habits we use, with prompts from our own threads:
Give it the whole app to run. Models are very good at writing tests their own code passes. Running the app end to end is the best way to know it works. Make that one command with
.agents/setup, services, and a skill. Tim hit a sidebar bug in our Mac app and asked:Repro it in preflight then we can look at fixing it.
It launched our Mac app, dragged the sidebar closed, and measured the bug before touching any code.
Tell it what done looks like, including the evidence it should deliver. Our bug fix prompts include:
Implement the simplest correct fix, verify it, then explain the bug, the fix, and the evidence that the fix resolves it.
That evidence is usually a failing test plus before and after videos.
Raise the bar when the check is too easy. If it verified against an emulator or a mock, send it to the real thing. When it checked an iPad fix in an emulator, Quinn replied:
test all of this in the Buildkite iOS preflight and use an iPad simulator
The simulator caught two bugs the emulator had missed.
Hand it the bigger task and go hands off. It holds up well over long threads. Tell it what evidence to deliver, then go do something else. You shouldn't babysit a capable model: if you have the patience to watch it work step by step, you're giving it too short a leash.
-