🏡


  1. September 30, 2026
    1. 🔗 backnotprop/plannotator v0.27.23 release

      chore: bump version to 0.27.23

    2. 🔗 earendil-works/pi v0.99.2 release

      New Features

      • MCP servers stay out of the way: servers with the default codemode exposure are no longer listed in the codemode description and no longer block the first prompt. They appear in a short system prompt section, and scripts find their tools with searchTools() and describeNamespace(). See Control tool exposure.
      • More MCP authentication options: oauth.clientName for servers that only accept known OAuth clients, and "auth": { "provider": "<provider>" } to authenticate HTTP servers with a provider's /login token. See Authenticate with OAuth.
      • Anthropic workload identity federation from the Anthropic SDK environment variables. See Use an API key from the environment.
      • /reload enables tools newly added to the defaultTools setting. See Tools.

      Added

      • Added a description field for MCP servers (pi mcp add --description), shown with the server in the system prompt and used to rank its tools in tool search, and a describeNamespace(name) codemode helper that returns a namespace's instructions and tool names. describeNamespace() and searchTools() accept a namespace as mcp__dev-radius, mcp__dev_radius, dev-radius, or dev_radius.
      • Added an oauth.clientName setting for MCP servers (pi mcp add --oauth-client-name) to change the client name sent during OAuth client registration, for servers that only accept known clients (#10226).
      • Added "auth": { "provider": "<provider>" } for HTTP MCP servers to send a provider's current /login token as the bearer token instead of using MCP OAuth. The token is read on every request, so provider refreshes apply. Only allowed in the global mcp.json and from extensions, and requires https except on loopback hosts.
      • Added Anthropic workload identity federation from the ANTHROPIC_FEDERATION_RULE_ID, ANTHROPIC_ORGANIZATION_ID, and ANTHROPIC_IDENTITY_TOKEN_FILE environment variables (see Providers) (#10177, #10242 by @philfreo).
      • /reload now enables tools newly added to the defaultTools setting. Tools removed from it stay enabled, tools turned off during the session stay off unless newly added, and --tools, --no-tools, and --no-builtin-tools still override the setting (#10245).

      Changed

      • MCP servers with the default codemode exposure no longer appear in the codemode description; scripts find them with searchTools(). codemode-deferred is now an alias for codemode. Use direct exposure for tools the model should see without searching (#10212).
      • The codemode description no longer includes deferred tools, tool counts, or MCP server instructions, so it no longer changes when MCP servers connect or change their tools. The tool_search description no longer lists the servers whose tools it can load, for the same reason. Servers are listed instead in an mcp_servers system prompt section with a one-line summary, updated at the start of each prompt; a changed section is appended to the conversation. Scripts read server instructions with describeNamespace() (#10212).
      • The first prompt no longer waits for MCP servers without direct tools. They connect in the background and are waited for when a codemode script names them, a script searches tools, or tool_search runs (#10212).

      Fixed

      • Fixed new sessions intermittently ignoring the saved default model, or warning that no models are available, when it belongs to an extension-registered native provider with a stored credential (#9962, #10190 by @davidbrai).
      • Fixed the /mcp sign-in URL not being clickable when it wraps across lines, by emitting it as a terminal hyperlink with a Cmd/Ctrl+click to open line like /login (#10186).
      • Fixed codemode image() accepting malformed base64 data or unsupported image types, which persisted an invalid image block that made every later provider request fail with HTTP 400 (#10215).
      • Fixed codemode failing to start its script worker from the standalone Windows executable (#10204).
      • Fixed prompt submission slowing down with session length, because resolving the session's model selection looked up the model catalog once per assistant message (#10198).
      • Fixed model lookups slowing down for providers with a refreshed pi.dev catalog, because merging remote catalog models took quadratic time.
      • Fixed the built-in-tool-renderer.ts and minimal-mode.ts extension examples removing the built-in tools' summaries and guidelines from the system prompt (#10072, #10193 by @christianklotz).
      • Fixed context overflow detection for Z.AI CN endpoint Prompt exceeds max length errors (#10208).
      • Fixed Anthropic requests failing when a tool schema uses keywords Anthropic strict tool use rejects, such as minimum/maximum; such tools are now sent non-strict (#9953).
      • Fixed provider retries firing immediately when a Retry-After header contains an unparseable date; they now use exponential backoff (#9571).
      • Fixed extension commands registered without a string name or handler crashing pi when typing /; the extension now fails to load with an error instead (#10054).
      • Fixed collapsed codemode and MCP tool results filling the screen when the output is one long line, such as minified JSON. Like bash output, the preview is now limited to wrapped lines instead of logical lines.
      • Fixed codemode.mode: "only" listing read, bash, edit, and write in the system prompt's tool list although requests only declare codemode (#10192).
      • Fixed codemode scripts calling the wrong MCP tool when two tool names differ only in - and _, such as read-file and read_file. Like in Codex, MCP tool and namespace names now replace - with _ (mcp__my-server__x is now mcp__my_server__x), colliding tools of a server all get a hash suffix, and server names that differ only in - and _ are rejected (#10239).
    3. 🔗 anthropics/claude-code v2.1.286 release

      What's changed

      • Added a count such as "2 of 5" to the permission prompt when several permission requests stack up
      • Added mouse support for the "N more" rows of lists in fullscreen mode: click one to jump to that end of the list, with hover and pressed states
      • Fixed several Claude Code processes and IDE extensions each opening a login browser when gcpAuthRefresh or awsAuthRefresh credentials expire
      • Fixed claude --resume and --continue sometimes losing every turn after a batch of parallel tool calls when the earlier session crashed or was killed
      • Fixed API 400 errors after a tool or hook returned an object, number or boolean instead of text, including in resumed sessions
      • Fixed cloud sessions with very large histories never waking up because the container was stopped while the transcript was still loading
      • Fixed the Claude apps gateway's spend meter pricing 1-hour prompt cache writes at the cheaper 5-minute rate, and counting only the first model call's input tokens on streamed turns that run a server-side tool such as web search
      • Fixed macOS sessions still showing "Not logged in" or "Login expired" after /login succeeds in another Claude Code window when a leftover ~/.claude/.credentials.json exists
      • Fixed every turn failing when the Anthropic API refuses the model your default or a model alias resolves to: Claude Code now retries once on the previous model of the same tier
      • Fixed Remote Control sessions (including claude remote-control) staying connected after your organization's policy turns Remote Control off; they now disconnect with a notice
      • Fixed refusal and --fallback-model retries failing when the fallback model can't run fast; they now run at standard speed, with a one-time notice in interactive sessions
      • Fixed headless sessions repeating the "MCP servers require authentication" reminder after a successful re-authentication when the MCP discovery cache is enabled
      • Fixed claude auth status reporting a Console sign-in's stored API key as claude.ai; it now reports api_key, and the VS Code extension treats that session as an API key session
      • Fixed /status listing an Anthropic profile beside an API key as if both were in effect; the profile is now marked as not in use
      • Fixed Claude not being told when a file attached to a message sent over Remote Control did not arrive, and a file sometimes getting only 10 seconds for its last download try
      • Fixed a Remote Control message that arrived while Claude Code was exiting being marked delivered and then never answered; it now stays queued for the session's next run
      • Fixed MCP error messages showing a credential's value when "Bearer" or "Basic" came before its key name
      • Fixed percent-encoded Bearer tokens being only partly masked in error messages
      • Fixed redacted logs and transcripts showing a secret whose key name has an invisible character inside, such as a zero-width space
      • Fixed logs and transcripts showing part of a URL password that contains punctuation such as ), quotes, ], & or a second @, or that runs past a / to a bracketed host such as [::1] in an ssh URL
      • Fixed the session transcript in the zip that /feedback saves to disk containing invalid JSON lines after secret redaction
      • Fixed MCP connectors listing no tools for up to a day after their server dropped the older MCP handshake
      • Fixed a repeat MCP sign-in request from Claude replacing the pending sign-in link, which could stop that link from working
      • Fixed /usage not crediting an MCP server for a tool call made while that server was still connecting or had only just connected, such as right after a restart
      • Fixed plugins enabled on claude.ai occasionally going missing from Claude Code for a session after a transient server error
      • Fixed a message typed into a running subagent showing twice in its transcript after the subagent read it
      • Fixed subagent hand-back messages showing a raw task id instead of the agent's name when the subagent had no registered name
      • Fixed foreground subagents sometimes missing the task-tracking tools (TaskCreate/Get/Update/List, TodoWrite) in sessions that have them enabled
      • Fixed subagents spawned with worktree isolation loading the project CLAUDE.md and its imports a second time from the worktree copy on their first file read
      • Fixed Workflow tool subagents being restarted from their original prompt when a connection stalled for a few minutes mid-response
      • Fixed /compact, /clear, and /rewind typed while viewing a background agent's or teammate's transcript silently acting on the main conversation: a dialog now names the target and asks first
      • Fixed background jobs showing done while waiting for your approval
      • Fixed the commit attribution reminder being re-sent inside tool output when a model fallback lasts only one turn
      • Fixed a click on the space between words of a collapsed row (such as "Thought for 4s") highlighting the row without expanding it in fullscreen mode
      • Fixed a row with no details, such as an action row with a long name, pushing every other row's details to the right in list screens
      • Fixed files with very long names not reaching a cloud session when attached to it
      • Fixed plugin errors for a marketplace Claude Code refuses to load: they now say why and how to fix it instead of "not found"
      • Fixed /plugin's Discover tab showing a marketplace name unquoted in its "Checking … for new plugins" line when its rows already show that name in quotes
      • Improved commit guidance: when your project or user skills include one named verify, Claude is now told to run it right before committing, except for docs-only and tests-only commits
      • Improved send now (ctrl+enter) in a subagent's view: it now moves the subagent's running command to the background so your message is read right away
      • Improved replies from background agents to your messages so they no longer open with a separate recap of what you said
      • Improved claude.ai artifact link reads: WebFetch now asks the same questions as the Artifact tool's read (no artifact prompt while the session's network access is on, one per artifact while it is off), and an auto-mode yes no longer counts where only you can answer
      • Improved fetch, skill, file read, sandbox network, Claude in Chrome, workflow script and notebook edit permission prompts to match the look of file edit prompts
      • Improved Bash, PowerShell and Monitor permission prompts to show the command between dashed lines, matching file edit prompts
      • Improved list scrollbars in fullscreen mode: in most lists the bar no longer shifts as the "N more" rows come and go, and it now has ↑/↓ arrows you can click or hold to scroll
      • Improved the external editor (Ctrl+G): editors that take a line number now open on the line your cursor is on in the prompt
      • Improved slash command suggestion responsiveness while typing when many skills or plugin commands are installed; command descriptions now match by word prefix
      • Improved the output style picker: it now opens on your current style instead of Default, with each style's description on the line under its name; number keys no longer pick a style
      • Improved /hooks: the closing line of a hook's detail screen now says "this hook" instead of "it"
      • Improved the model fallback notice and the autocompact-thrashing error to say when a fallback dropped the context window from 1M to 200K tokens
      • Improved responsiveness of SDK and -p sessions when a host re-sends an MCP server enable for a server that is already connected
      • Improved the protocol page a Claude apps gateway serves at /protocol: it now says not to reject unknown input and matches what Claude Code sends today
      • Changed prompts sent while nothing is running or queued to show in the normal text color right away instead of gray
      • Changed how failed API requests are retried: one limit now covers a whole model call, so with the default retry settings a failing call sends at most 14 requests
      • Changed --bare to connect only the MCP servers named on the command line, send the model no system reminders, and start no background tasks; under --bare, a shell command that reaches its timeout now stops instead of moving to the background
      • Changed the send-now key (ctrl+enter) to move a skill's own shell command to the background instead of ending it
      • Changed the WebFetch error for a rate-limited domain safety check to tell Claude not to retry it in a loop
      • Changed plugin installs to refuse npm sources that are git repositories or folders, and to install plugin dependencies only from registry packages
      • Changed list screens (/artifacts, /mcp, /skills, /hooks and others) to always line up each row's details in one column after the names
      • Changed the overflow rows of lists to read "↑ N more" / "↓ N more" instead of "N more above" / "N more below"
      • Changed /hooks to open on one list of your configured hooks grouped by event, so viewing a hook takes one Enter instead of three
      • Changed the theme picker to a scrolling list that fits your terminal instead of pushing the preview off screen; number keys no longer pick a theme
      • Changed /exit's Remove worktree to run after Claude Code stops the servers and shells it started there, which on Windows could keep the folder from being deleted
      • Changed the claude-api skill's Managed Agents examples to create environments with limited networking
      • Removed the browser link from /ultrareview and claude ultrareview output
      • Windows: Fixed claude --bg and the agents view refusing a folder that claude already trusts when its trust record was saved with different letter case
      • [VSCode] Added bookmarks: save Claude's responses and keep them in view in a Bookmarks side panel
      • [VSCode] Added the questions Claude asks and your answers to the conversation: after you answer a question card, a Questions row shows each question with your picks
      • [VSCode] Added option previews to question cards in the chat panel: the highlighted choice's mockup or snippet shows beside or under the options
      • [VSCode] Added rows under a message that open to the terminal output, browser tab, browser instructions and selected code sent with it
      • [VSCode] Fixed a second copy of a conversation opening in a tab when it was already open in the side bar; the side bar now switches to it
      • [VSCode] Fixed settings dialogs reporting a failed save, without re-checking, when Claude Code printed more than 1 MB of output
      • [VSCode] Fixed an endless "Teleporting session…" spinner when the extension stops responding
      • [VSCode] Improved the Manage plugins dialog: it says when a turned-off plugin is still on because of other settings, and explains a plugin folder clash
      • [VSCode] Changed Stop and Escape to end only the current turn; background agents keep running and can be stopped one by one from the agent map
      • [VSCode] Changed the "✻ Claude Code" status bar item to show in every window, so you can open Claude when no file is open
      • [Cloud sessions] Fixed an answered question card or approved tool call getting no reply when the session had gone idle after Claude sent a message
      • [Cloud sessions] Fixed clearing an organization environment's setup script in admin settings leaving new cloud sessions still running the old script
      • [Cloud sessions] Fixed the Runner actions menu on the self-hosted environments admin page closing on its own a few seconds after it opened
      • [Cloud sessions] Fixed routine runs whose cloud session never started showing as Succeeded in the Runs pane, the routine's page and the sidebar; they now show as Failed
      • [Cloud sessions] Fixed clicking an audio or video file in a cloud session's Outputs card opening an empty file search instead of playing the file
      • [Cloud sessions] Changed a routine's page to read "Due" with the scheduled time, instead of a next run time in the past, when a scheduled run is late and hasn't started
      • [Claude Tag] Added an Add channel button to Claude Tag's spend limits page in admin settings, so a limit can be set on any channel, including a private one, from its channel ID or Slack link
      • [Claude Tag] Fixed memory recall finding nothing in organizations that cannot use the default Sonnet model
      • [Claude Tag] Fixed a Slack channel losing its Claude settings when an Enterprise Grid admin moves it to another workspace and the first post afterward doesn't mention Claude
      • [Claude Tag] Fixed Claude in a Slack thread occasionally starting over on a fresh machine, losing work it hadn't pushed, when your reply answered a question it had just asked
      • [Claude Tag] Fixed public channel names on Claude Tag's spend limits page in admin settings showing as raw Slack IDs in larger organizations
      • [Claude Tag] Improved the titles of sessions started from Slack as shown on claude.ai: they now read as the words you typed, without Slack user IDs or escape codes
    4. 🔗 crmne/spotifast Spotifast v0.11.1 release

      Spotifast 0.11.1 shows emoji in colour everywhere, comes as an AppImage, selects songs with the arrow keys, and wears a refreshed icon. It also fixes local playback on Windows on ARM, the volume jumping back after a reconnect, and a dozen smaller problems reported since 0.11.0.

      Download Spotifast: Mac · Windows · Windows ARM · Linux · Linux ARM · Flatpak · AppImage · AppImage ARM

      Spotifast's icon before and after: the play triangle is larger and centred,
on a polished
disc

      The icon before and after.

      New

      • Emoji in colour, everywhere. Names, lyrics, menus, tooltips and text fields such as the search box now draw emoji in colour, in the system's own emoji font. By @crmne.
      • Spotifast as an AppImage. Each release now includes spotifast-X.Y.Z-x86_64.AppImage and spotifast-X.Y.Z-aarch64.AppImage: one file, make it executable and run it, no installation. By @crmne.
      • Select songs with the arrow keys, remove them with Delete. In a playlist, album or Liked Songs, up and down now select the song they land on, and Shift extends the selection. In a playlist you can edit, Delete (and Backspace on macOS) removes the selected songs. By @shrinikets. (#585)
      • Remove from this playlist on the playing song. Right-clicking the song in the player bar now offers the removal while it plays from a playlist you can edit. By @TacticalDeux. (#592)
      • A refreshed icon. The play triangle is larger, has softly rounded corners and sits in the centre of the disc, and the disc has a polished rim and a lit face. By @crmne.

      Fixed

      • Local playback works on Windows on ARM. Turning it on answered "Premium account required" on ARM computers, because Spotify did not recognise the platform Spotifast announced. By @crmne; thanks @watermarkhu. (#123)
      • The volume you set survives a reconnect. When the Spotify session dropped and came back, the volume returned to the level Spotifast started with: music turned down to 5% could come back at 80%. A volume set before this computer became the active device is kept too. By @queueeee. (#589)
      • The window drags from the top of every panel with the custom title bar. On Windows, grabbing the top of the window over the sidebar, the queue or the lyrics did nothing. By @Outfluencer. (#590)
      • Clearing the search stays on the page you are on. The clear button opened the search page. By @theworker02; thanks @D3SOX. (#609, #596)
      • A single song starts on a phone or speaker. Playing or resuming a one-song context on another device failed with "Non supported context uri". By @Magniquick. (#605)
      • Spotify's own mixes line up. A playlist whose songs carry no added dates no longer reserves an empty Date added column that pushed the album names out from under their heading. By @Magniquick. (#607)
      • The visualizer's tooltip stays off the controls. "Click to change the visualizer" followed the pointer onto the play button and the seek bar and covered their tooltips. By @Magniquick. (#606)
      • A settings file saved by Windows PowerShell or Notepad is read. Such a file starts with a byte order mark; Spotifast rejected it, fell back to the defaults and then saved them over your settings. By @kevin9327. (#617)
      • A skin unpacked with Windows' Extract All appears in the list. Its files sit one folder deeper, which Spotifast could read but did not list. By @kevin9327. (#618)
      • The MilkDrop sliders match the other settings. Time per preset and Frame rate now have the same shape, font and colours as the controls around them. By @nos1609. (#600)
      • The cover stays put while full-screen lyrics load. It moved aside for the lyrics and back again when a song had none; it now waits in the middle, with "Loading…" beneath it, and moves only once there are words to show. By @crmne.

      Known limitations

      • The AppImage bundles no libraries: it needs glibc 2.39 or newer and your desktop's own libraries, like the DEB and RPM. It does not update itself; download the new file when Spotifast says a release is out. Running it needs FUSE, or --appimage-extract-and-run.

      Thanks

      @shrinikets, @TacticalDeux, @queueeee, @Outfluencer, @theworker02, @Magniquick, @kevin9327, @nos1609, @watermarkhu, @D3SOX, and everyone who reported problems since 0.11.0.

      Full changelog : v0.11.0...v0.11.1

    5. 🔗 OmniNull/OmniWM OmniWM v0.7.4 release

      What's New Since 0.7.3

      Name your windows, place the workspace bar on any screen edge, and choose OmniWM’s interface language. OmniWM 0.7.4 also adds Niri controls and moves blocking focus and window queries off the main thread.

      Window marks and language support

      • Find windows by name. Give live windows marks from the command palette or omniwmctl. Mark matches appear first in window search; focus a marked window across workspaces or summon it beside the focused tiled window. Marks last until the window closes or OmniWM quits.
      • Assign mark shortcuts. Set Mark and Remove Mark are available in Settings > Hotkeys and start unassigned. The palette also provides buttons and shortcuts for marking its selected window.
      • Bring recent windows forward. With an empty search, Windows mode lists the focused and recently focused windows first. Shift + Enter can move the selected window into an empty current workspace, including floating windows; floating windows cannot be summoned to the right.
      • Choose the app language. Settings > General now lets OmniWM use a different interface language from macOS. The default follows macOS, and changes apply the next time OmniWM starts.

      Workspace and layout behavior

      • Place the workspace bar on any edge. Choose top, bottom, left, or right placement, with matching layout space, menus, and previews. Hidden Bar follows bar size changes and reveals concealed icons when the global workspace bar is disabled.
      • Use shortcuts above workspace nine. Assign Switch, Move, and Move Column shortcuts to higher-numbered workspaces, including dynamic workspaces.
      • Skip empty workspaces consistently. When Hide Empty Workspaces is enabled, relative navigation, workspace swipes, and Overview skip inactive empty workspaces. Direct workspace selection remains available.
      • Control Niri sizing and gaps. Grow and shrink increments are now configurable from 1% to 100%, defaulting to 5% instead of the previous 10%. A new Gaps at Screen Edges option lets inner gaps appear only between windows.
      • Keep Niri positioning predictable. Explicitly centered first and last columns stay centered during ordinary relayouts. A brief pill shows the column’s tabbed state, including when it contains only one window.
      • Avoid Dwindle fullscreen overlap. Arriving tiled windows, including restored windows, end layout fullscreen before being added.

      Responsiveness and feature controls

      • Less blocking work during keyboard navigation. Niri keyboard and wheel focus now front apps off the main thread and target the chosen window directly, avoiding an initial activation of the app's previously active window. Layout passes skip a roughly 9 ms Screen Recording permission check once the wallpaper preview is captured, and Hidden Bar no longer rescans running apps on every app switch.
      • Measured keyboard gains. In captures with matching display-link timing, estimated missed animation frames per focus-changing keypress fell from 2.5 to 0.125. Median hotkey handling fell from 6.5 ms to 0.44 ms. Median time to the first animation callback fell from 18.7 to 9.0 ms across apps and from 10.9 to 7.1 ms within the same app.
      • Slow window queries move off the main thread. Ordinary frame-change WindowServer queries, including one recorded half-second wait, now run on a worker. New-window lookups and window lifecycle and rule metadata queries also run off the main thread. Closed windows release their cached previews.
      • Choose which features run. Overview and workspace bar hover previews can be disabled independently. Disabled features release their panels, observers, previews, and retained resources; the global workspace bar switch now takes precedence over monitor overrides.
      • Updated the embedded GhosttyKit terminal framework to 1.3.2-main+76895d97b (upstream changes).
      • Diagnostic captures provide more detail about focus, parking, frame reads, and preview activity.

      Performance figures come from instrumented Release development captures on one Mac. They measure internal handling and animation callback timing; screen presentation latency was not measured.

      Upgrade notes

      • settings.toml upgrades to schema 4 on first launch. The original is kept as settings.toml.pre-v4 (settings.toml.pre-v4.1 if needed); the rewritten file removes comments. Restore the backup before downgrading to 0.7.3.
      • Direct IPC clients must use protocol 17. The omniwmctl included in this release already uses it.

      Thanks

      Thank you to:

      Full changelog: v0.7.3…v0.7.4

      Website and documentation · Installation guide

      Release Integrity

      The app is signed, notarized, and stapled. SHA-256 hashes:

      • OmniWM-v0.7.4.zip: 1f1485e1d166831697e8cedb9c117577f5949fe00deeae8572bed6d6c58eb4e2
      • GhosttyKit.xcframework-v0.7.4.zip: 59e16cccd81ed87fde098bc40122154e640b950ba3b7b997879c786730b88c8e
    6. 🔗 crosspoint-reader/crosspoint-reader 1.6.5 release

      Summary

      Library view

      Recent Books has grown into a powerful way to browse your entire collection on your SD card. Sort by recently added, title, or author, or just search for the book you want.

      TTF support

      Devices with external RAM (X4Pro, Sticky, X4C, PaperMono) can now use TrueType (.ttf) fonts directly — no need to convert them to CrossPoint's .cpfont format first. Drop your fonts into /fonts/ or /.fonts/ on the SD card, restart the reader, and pick them from the text settings. If a family has separate regular, bold , italic , and bold-italic files, put them together in one family folder.

      .cpfont fonts still work everywhere, including on devices without external RAM. OpenType (.otf) fonts are also supported, though compatibility may vary.

      Cover Grid home theme

      The new Cover Grid theme displays your recent books as a grid of covers on the home screen. This theme is not available on x3 and the original x4.

      More control over reading

      New word and character spacing controls let you adjust how tightly text sits on the page, alongside the existing line-spacing settings. Footnote navigation now highlights references directly on the reading page and and bookmarks can finally be renamed.

      Touch controls and navigation

      Touch devices gain configurable tap and swipe gestures for turning pages. Headers now have back buttons, and fixes improve scrolling in reader lists and settings. X4pro also gain configurable shortcuts for Home button.

      X4 Classic support

      The ESP32-S3-based X4 Classic is now officially supported. Get one now

      Sleep Screens

      Sleep screens across all devices now show drastically higher quality grays. On the X4 and certain Pro display variants they push the quality to the max. Transparent pngs got a nice quality bump as well.

      Clock & About

      You can now view the clock on the home screen! In addition, the old utc offset picker is gone and is replaced with auto DST and a list of cities to pick from when setting the time. This setting has moved to the more obvious system settings. You'll also find a new About option in the system settings as well which is useful for identifying display controllers during debugging.

      Everything else

      Files can be renamed straight from the file browser. Portuguese gets hyphenation support, Korean text justification stretches spaces between words rather than within them, and the keyboard gains an Arabic layout.

      KOSync now sends more precise EPUB reading positions. Fixes also address EPUB lists, hidden content, chapter-position displays, and end-of-book navigation.

      This release reduces font and EPUB memory pressure, fixes USB drive disconnection, improves web file-transfer safety, and prevents EPUBs that fail to render their first page from reopening automatically after wake.


      What's Changed

      New Contributors

      Full Changelog : 1.6.0...1.6.5


      Downloads

      crosspoint-1.6.5-papermono.bin
      crosspoint-1.6.5-sticky.bin
      crosspoint-1.6.5-x3-x4.bin
      crosspoint-1.6.5-x4c.bin
      crosspoint-1.6.5-x4pro.bin

    7. 🔗 modem-dev/hunk v0.23.0 release

      What's Changed

      • fix(ui): open help and menus without first-use suspension by @elucid in #1079
      • test(pty): synchronize inputs on committed screen state by @elucid in #1083
      • perf(highlight): load Shiki WASM bytes without base64 decoding by @elucid in #1078
      • docs(release): publish v0.22.0 notes by @benvinegar in #1087
      • perf(test): reduce fixture process overhead by @benvinegar in #1092
      • feat(ui): host-owned status line with inline prompts, exposed to extensions by @elucid in #1095
      • feat(search): / searches diff content via a bundled search extension by @elucid in #1096
      • Make daemon/client build skew visible and resolvable with hunk daemon status/restart by @elucid in #1099
      • feat(config): make wheel scrolling configurable by @benvinegar in #853
      • feat(extensions): add host-owned file-view syntax highlighting by @benvinegar in #1053
      • feat(ui): add support line number in "Open file in editor" for zed by @seiichi1101 in #1073
      • Add hunk-triage to hunk-extension marketplace by @benvinegar in #1113
      • test(pty): speed up integration synchronization by @benvinegar in #1098
      • fix(nix): install hunkdiff alias so hunk patch works by @schickling-assistant in #1108
      • fix(test): compare install VM paths in one spelling so symlinked tmpdirs pass by @shashank-100 in #1120
      • fix(cli): resolve bundled skills by proximity, not candidate order by @shashank-100 in #1116
      • docs(release): gate Homebrew availability claims by @benvinegar in #1088
      • fix(ui): register one terminal blur listener per diff pane by @speedsharmaai in #1111
      • fix(nix): use stdenv.hostPlatform.isLinux by @patrickhaahr in #1101
      • Add hunk-starlark to extension marketplace by @HackAttack in #1124
      • feat(theme): add a default terminal theme that follows terminal colors by @dhh in #1128
      • feat(extensions): bundle GitHub review commands by @benvinegar in #1129
      • chore(release): prepare v0.23.0 by @benvinegar in #1133

      New Contributors

      Full Changelog : v0.22.0...v0.23.0

    8. 🔗 exe.dev Show and Tell: Shelley rss

      Shelley learned a new trick recently: instead of typing a prompt, you can record both your speech and your window (or screen). Once you stop recording, Shelley puts together the transcript and the matching video, and hands that off as the prompt to the agent.

      Shelley's message box with the video camera record button circled in red,
between the attachment paperclip and the send
button

      You have to see this to believe it. Shelley had a bug where if you sent off a “Compact and Send” and navigated to a different conversation, the navigation would revert back to the original one. I could have described it, but after it happened to me, I started a new conversation, and took the following recording:

      Your browser does not support the video tag.

      The recording is not edited except my shortcuts and URL bar are cut off, some commentary is added below, and it’s sped up. It’s painful to watch (especially for me!): all of those ums, and figuring out what to click to reproduce, and so on. And yet, the next message from me in that conversation was “push.” The agent was able to reproduce based on what it saw and fix the bug, in one-shot, and I typed nothing.

      Give it a ~~shot~~ screen recording!

    9. 🔗 Cryptography & Security Newsletter Get Ready for Seven-Day Certificates rss

      With the groundwork for the post-quantum migration of key establishment behind us, browser vendors are increasing their efforts on the remaining parts: post-quantum authentication. (If you haven’t been following the events surrounding post-quantum migration, our newsletter from two months ago has a quick recap. Read it before continuing here.)

    10. 🔗 smol-machines/smolvm smolvm v1.22.0 release

      What's Changed

      Full Changelog : v1.21.1...v1.22.0

    11. 🔗 smol-machines/smolvm smolvm v1.21.1 release

      What's Changed

      • Stream macOS checkpoints sparsely so saves scale with memory in use by @BinSquare in #1484
      • Reword the README description, add an SDK example, and explain shared responsibility by @BinSquare in #1483
      • Bump the workspace to 1.21.1 by @BinSquare in #1485

      Full Changelog : v1.21.0...v1.21.1

  2. September 29, 2026
    1. 🔗 smol-machines/smolvm smolvm v1.21.0 release

      What's Changed

      • Let machine update change a stopped machine's egress allow list by @BinSquare in #1471
      • Bump the Nix package to 1.20.2 by @BinSquare in #1472
      • Support Windows memory checkpoints and frozen branches by @BinSquare in #1474
      • Restore a checkpoint's layered disk as a copy-on-write top instead of copying its top layer by @BinSquare in #1475
      • Let a stopping server finish in-flight requests for a configurable grace by @BinSquare in #1473
      • Cut machine exec latency: wake on exit, stop crun re-cloning itself, fork less in the wrapper by @LoganGrasby in #1479
      • Speed up cold checkpoint restore by hashing during extraction by @BinSquare in #1480
      • Let embedders create detached machines and list every machine by @BinSquare in #1478
      • Key a restored machine's VMM uid on its own data dir by @BinSquare in #1476
      • Treat an empty machine start body as no body by @BinSquare in #1481
      • Start registry-image machines on a shared seed of their image by @LoganGrasby in #1477
      • Bump the workspace to 1.21.0 by @BinSquare in #1482

      Full Changelog : v1.20.2...v1.21.0

    2. 🔗 imfing/hextra v0.13.0 release

      Hextra v0.13.0 brings a brand-new search experience that works smoothly on both mobile and desktop, an image gallery shortcode with a lightbox, foldable alerts, remote icon packs, and better Jupyter notebook rendering, plus a round of search improvements and bug fixes.

      For the full release notes and upgrade guide, please visit:
      https://imfing.github.io/hextra/blog/v0.13/

      New Features

      • New Search Experience : search is now a command palette dialog that works smoothly on both mobile and desktop. Open it with ⌘K / Ctrl+K or / on desktop, or tap the search icon on mobile. Results show breadcrumbs, and the whole dialog works from the keyboard.

      • Gallery Shortcode : image galleries with a PhotoSwipe lightbox and grid, carousel, and mosaic layouts
      • Foldable Alerts : + makes an alert foldable and expanded by default, - starts it collapsed, with an optional custom title (> [!TIP]- Title), following Obsidian callout syntax
      • Remote Icon Packs : use Lucide, Tabler Icons, and Simple Icons with a provider prefix, e.g. {{< icon "lucide:rocket" >}}, anywhere Hextra accepts an icon name
      • Improved Jupyter Rendering : many more output types (error tracebacks, stderr, SVG, Markdown, LaTeX, JSON, raw cells, attachments), plus In [N]:/Out[N]: prompts via prompts=true
      • Custom Page Sections : add content before and after each page and its body with the page-begin, content-begin, content-end, and page-end partials
      • Search Improvements : cleaner results without stray HTML tags, correct links to headings with inline markup, and a fingerprinted search index in production

      Search CSS Changes

      The inline search box has been replaced by a dialog, so the hextra-search- wrapper class no longer exists. If you've styled the search UI with custom CSS, target the new hextra-search-trigger and hextra-search-dialog classes instead.

      What's Changed

      • chore: update Hugo language config keys by @imfing in #993
      • feat: support remote icon packs by @imfing in #995
      • feat: add gallery shortcode with PhotoSwipe lightbox by @imfing in #996
      • feat: improve Jupyter notebook shortcode rendering by @imfing in #1001
      • feat: add foldable alert support by @imfing in #1002
      • fix: resolve relative markdown links in render hook by @imfing in #998
      • feat(search): replace inline search with command palette dialog by @imfing in #1003
      • feat: add custom page section partials by @imfing in #1004
      • fix(search): fingerprint search data in production by @imfing in #1014
      • fix: support asciidoctor page headings by @imfing in #1005
      • chore(deps): bump postcss from 8.5.14 to 8.5.25 by @dependabot[bot] in #1018
      • fix(search): handle inline markup in heading fragments by @imfing in #1015
      • feat(search): Drop HTML tags from search result by @bombsimon in #1020
      • docs(showcase): add Roled documentation by @muhammadkholidb in #1031
      • chore(deps): bump nanoid to 3.3.19 by @imfing in #1033
      • fix(katex): select compatible CSS for Hugo renderer by @imfing in #1029
      • fix(mermaid): correct diagram scaling with reduced motion by @yuri1969 in #1030
      • fix(mermaid): render diagrams lazily when revealed in hidden containers by @imfing in #1034
      • fix(sidebar): fall back to page tree in mobile menu without page-backed entries by @imfing in #1035
      • fix: close gallery shortcode comments correctly by @avighnac in #1041
      • chore: update link to use safeURL function by @farmacia-cambie in #1037
      • fix: missing spaces between attributes, and unclosed hero tags when nested by @hoshsadiq in #1039
      • test: fix render-link expectation after attribute spacing fix by @imfing in #1042

      New Contributors

      Full Changelog : v0.12.3...v0.13.0

    3. 🔗 anthropics/claude-code v2.1.285 release

      What's changed

      • Added CLAUDE_CODE_DISABLE_WEB_FETCH environment variable to turn off the WebFetch tool
      • Added claude --desktop to open the Claude desktop app on the current directory, or on a session with --continue / --resume <id>
      • Added claude plugin configure <plugin> to show a plugin's options and which are unset, or save new values read from stdin with --values-stdin
      • Added <server>.<key>=<value> to claude plugin install --config, so a bundled .mcpb MCP server's own settings can be set at install time and it starts without visiting /plugin → Configure
      • Added allowedProviders managed setting to limit which API providers a machine may use (Anthropic API, a custom endpoint, Bedrock, Mantle, Vertex AI, Foundry, Claude Platform on AWS, or a Cloud gateway)
      • Added CLAUDE_CODE_NONSTREAMING_TIMEOUT_RETRIES environment variable to cap re-sends of a non-streaming fallback request that timed out
      • Fixed claude -p with CLAUDE_CODE_FORK_SUBAGENT=1: a subagent's own Agent call now runs in the foreground, so the subagent gets the child's result
      • Fixed plugin and marketplace installs and updates over SSH ignoring the ssh program set in GIT_SSH or in your git config's core.sshCommand
      • Fixed Claude Code refusing to start when the OS denies reading the managed settings file; it now warns and starts without that file's policies. Other read errors and unparseable files stop every session
      • Fixed cloud sessions that restarted after their conversation was compacted refusing the next update to an artifact the session had already read or published
      • Fixed claude plugin disable and enable with a full name@marketplace id changing a settings entry in another letter case instead of the installed plugin's own
      • Fixed files attached to a message sent over Remote Control being left out after a single failed download; a network error, timeout or server error is now retried up to twice
      • Fixed switching models mid-session with a set_model request (such as the Agent SDK's setModel) leaving the new model on the built-in output-token limit and auto-compact window until restart
      • Fixed redacted logs and transcripts showing part of a URL password that contains @, or all of it when the URL writes its @ as %40
      • Fixed SSH passphrase and new-host prompts from worktree and /teleport fetches taking over the terminal; these fetches now fail fast instead of asking
      • Fixed switching off an MCP server added mid-session in SDK and -p sessions leaving its tools available
      • Fixed claude -p --permission-prompt-tool: a background subagent's permission request now goes to the prompt tool instead of being auto-denied
      • Fixed claude mcp list and claude mcp get, and the not-found error of claude mcp remove, login and logout, printing line breaks and terminal escape sequences from MCP server names and values
      • Fixed sandbox auto-allow asking for approval on every run of many inline scripts (python3 -c, node -e) just because they contain =
      • Fixed fork subagents not keeping the session's plan mode or dontAsk mode: a fork now runs under its parent's permission mode and cannot exit plan mode
      • Fixed claude remote-control --help saying --[no-]chrome defaults to the machine's /chrome setting; spawned sessions keep Claude in Chrome off unless --chrome is passed
      • Fixed background subagents in auto mode prompting a second, redundant reply after each report
      • Fixed cloud session creation and /remote-env reading only the newest 20 of an account's environments
      • Fixed Remote Control marking a message as read as soon as it arrived instead of when Claude started on it, and losing a message still queued when the terminal quit (it now arrives on the next resume)
      • Fixed installing a plugin with claude plugin install or /plugin putting it into an installed plugin's cache or data folder when their ids differ only in ., -, @ or (macOS, Windows) capitals; the install is now refused
      • Fixed hooks and SDK permission callbacks seeing a missing or outdated plan on ExitPlanMode when the plan was written in the same response
      • Fixed the first reply in cloud sessions arriving tens of milliseconds late, a regression in 2.1.283
      • Fixed sessions that authenticate with ANTHROPIC_AUTH_TOKEN against the Anthropic API never loading the organization's policy
      • Fixed a failed agent(), parallel() or pipeline() call that a workflow script awaits later, or not at all, being treated as an unhandled promise rejection, which could end a background session
      • Fixed synchronous hooks hanging Claude Code while a background process the hook started (for example some-daemon &) kept its output open; the hook now finishes shortly after its own process exits
      • Fixed WebFetch reporting a rate-limited domain safety check as a network or enterprise policy block
      • Fixed the fullscreen ctrl+o transcript freezing briefly when opened on turns with hundreds of file reads or searches; tool calls still running when the transcript opens now show their results when they finish
      • Fixed Amazon Bedrock mid-stream modelTimeoutException and serviceUnavailableException errors showing a raw JSON body instead of the error message
      • Fixed /autofix-pr and /schedule saying the Claude GitHub App is not installed on a repository whose install status had not been checked yet
      • Fixed dismissing a row (x) in /artifacts unlinking its file from the artifact, so publishing the same file again created a new artifact instead of updating it
      • Fixed Artifact tool publishes after a conversation rewind (Esc Esc) overwriting a file's newer content that Claude had read only in the rewound turns; the publish is now refused until Claude re-reads the file
      • Fixed an Artifact allow rule ("don't ask again") letting the Artifact tool publish a file outside the working directories without asking; add the file's folder with --add-dir for the rule to cover it
      • Fixed /cost and SDK modelUsage reporting a turn under the wrong model when the server answered a refusal with a different fallback model than the client expected
      • Fixed the Artifact tool so that publishing a page no longer lets Claude overwrite its source file without re-reading it when Claude's earlier read was cut short or the file had changed since
      • Fixed auto mode skipping its classifier for Artifact tool asset uploads and reads of someone else's artifact when you had approved that artifact earlier in another permission mode
      • Fixed a misleading "core.worktree is set" error from /ultrareview when the project folder briefly could not be read
      • Fixed /ultrareview on macOS and Linux failing to upload the working tree from a git worktree whose per-worktree config sets core.longpaths
      • Fixed the Artifact tool sometimes reporting a publish as a conflict with another session after retrying a temporary server error, when the first attempt had actually succeeded
      • Fixed /ultrareview uploads including uncommitted changes to credential files whose name has a colon before the extension, such as server:8443.key
      • Fixed a rare auth failure when two sessions recover a login refresh lock left by a crashed process at the same time
      • Fixed the PowerShell tool's permission check skipping deny and ask rules, and caching that failure for later checks, when its command parser failed to start (for example when the machine was out of memory)
      • Fixed /ultrareview uploads on macOS and Linux running slowly on some unusual file names, and their credential-file check missing file or folder names with many backup or editor marks
      • Fixed a cancelled shell command or hook still starting, and running to its end, when the cancel arrived while it was being set up
      • Fixed vim mode: after editing in the external editor (Ctrl+G), x or r in NORMAL mode no longer breaks a pasted-text placeholder at the end of the prompt
      • Fixed responses blocked by the API's output content filter being re-sent and retried, sometimes for minutes, instead of showing the filter's error right away
      • Fixed plugins silently skipping a bundled .mcpb MCP server that still needs configuration: /plugin, the install message and claude plugin install now say so and point to Configure
      • Fixed compacting or resuming a session failing, opening without its history, or crashing when its saved transcript holds a compaction marker or loop wakeup entry with missing or malformed fields
      • WSL: Fixed /ultrareview refusing to upload a checkout on a Linux volume when a changed file's name has a colon or ends in a dot or space
      • Fixed CLAUDE_CODE_RESUME_INTERRUPTED_TURN re-running a turn that had ended at --max-turns
      • Fixed sign-in that could wait forever after the browser showed success
      • Windows: Fixed /ultrareview uploading a linked worktree of a repository rooted at your home folder in some cases
      • Fixed cloud sessions reporting the uploads folder as missing before any file had been uploaded
      • Fixed a reply sent from claude agents to a background session waiting on a permission prompt sometimes approving the pending command
      • Fixed claude attach, logs, stop, respawn and rm starting a new session with the command name as its prompt when options came before it, such as from a shell alias
      • Fixed claude mcp list leaving out WebSocket (ws) MCP servers; each is now listed with its URL and health status
      • Fixed claude mcp get showing no Type, Command, Args, or Environment for stdio servers whose config entry omits the type field
      • Fixed .claude/settings.local.json allow rules being held back outside a git repository when git's trace2 output is configured
      • Fixed the /claude-api eval runner scaffold and report builder writing through a symlink or hard link planted at an output file
      • Fixed the /claude-api eval runner scaffold counting responses cut off at max_tokens in the score averages; they are now marked truncated and counted separately
      • Fixed &nbsp; showing as literal text in the terminal when a reply uses it to indent text, such as row labels in a markdown table
      • Fixed a brief freeze (up to a second) partway through long sessions outside fullscreen mode, which came back after /clear or /compact
      • Fixed an approved Edit never going through when its target is a device, such as a file symlinked to /dev/null, and the approval came from the IDE diff view or changed the edit
      • Fixed a failing API request being retried up to 21 times when streaming kept failing; the non-streaming fallback now shares the request's retry budget instead of getting a fresh set of retries
      • Improved Claude in Chrome: the native host now reports your computer's name, so connected browsers can be labeled by computer instead of "Browser 1" / "Browser 2"
      • Improved Bedrock and Vertex AI sessions to switch to an older available model of the same tier, instead of failing, when an admin removes access to the default model; session titles and summaries now fall back with it
      • Improved plugin marketplace errors to name why a git address is refused instead of citing enterprise policy
      • Improved validation of git URLs for plugins, marketplaces and the current repository's remote
      • Improved Artifact tool results: they now suggest publishing in the same step as writing or editing the page, which can save a round trip
      • Improved Remote Control: a /btw side question asked of a session hosted by an app such as Claude Desktop now sees the turn in progress, not only the last finished one
      • Improved pictures Claude sends as BMP, HEIC, HEIF, AVIF or TIFF files: the Claude apps now show a preview where Claude Code can convert them
      • Improved subagents in auto mode: a subagent's run now ends as soon as it hands its report back to its caller, instead of taking extra turns that reach no one
      • Improved Bedrock and Vertex start-up model checks: models your account cannot use are now remembered for up to a day instead of being re-checked on every launch
      • Improved Artifact tool publish results to use fewer tokens: the note on updating an artifact is shorter, and where to find your artifacts is no longer repeated after every publish
      • Improved SDK liveness during a non-streaming fallback request: with partial messages on, a ping stream event is now sent every 30 seconds on the Anthropic API, Claude Platform on AWS and gateways
      • Improved /resume and claude --resume on a session that is running in the background: they now open that session instead of refusing, and a prompt given with claude --resume <id> "prompt" is sent to it as its next turn
      • Improved per-turn performance when many permission deny rules and MCP tools are configured
      • Improved responsiveness when leaving the ctrl+o transcript view in long sessions when fullscreen rendering is off
      • Improved Bedrock, Vertex and Mantle start-up model checks to send the same User-Agent, x-app and session ID headers as regular requests
      • Changed MCP tools so a tool that sets its own _meta['anthropic/alwaysLoad'] to false stays deferred when its --mcp-config, Agent SDK or plugin server is set to alwaysLoad
      • Changed background Bash and PowerShell commands to stop after a time limit (their timeout with run_in_background, default 30 min, max 2 h); Claude is notified when one is stopped
      • Changed Code Review's pull request reviews and /ultrareview to run when disableWorkflows is on, unless the machine running the review has it set by its own administrator (MDM or the managed-settings file)
      • Changed sessions behind a custom ANTHROPIC_BASE_URL to use the 1M context window of models that have one (Opus 4.7+, Sonnet 5+, Fable); run /autocompact 200k if your gateway stops at 200K
      • Changed Team and Enterprise sessions, and sessions whose sign-in plan Claude Code can't determine, to withhold WebFetch until the organization policy loads if it couldn't be loaded at startup
      • Changed /memory so that Auto-memory can no longer be turned on from a background session or from a session one of Claude Code's own tools started; turning it off there still works
      • Changed the one-time offer to make auto mode your default permission mode to also show on third-party providers and with telemetry off, when your user settings default to another mode
      • Changed claude -p and Python Agent SDK sessions on third-party providers or with telemetry off to start in auto mode when no permission mode is configured, like interactive sessions; --permission-mode still overrides it
      • Changed Bedrock, Mantle and Claude Platform on AWS requests to a base URL with a non-default port to include the port in the SigV4-signed Host header
      • Changed the MCP server name widgets to be reserved in cloud sessions and on self-hosted runners: your own server under it, or a close spelling such as widgets_, no longer loads, so rename it
      • Changed /ultrareview on macOS and Linux to leave symbolic refs out when uploading a local checkout; a checkout whose current branch is a symbolic ref is now refused with an explanation
      • Windows: Changed project and local settings env to no longer set ALLUSERSPROFILE, SystemDrive, or the CommonProgramFiles variables; set them in user or managed settings instead
      • Changed /tasks to fold background work Claude Code runs for itself under one "System tasks" row; press Enter on it to show those tasks
      • Changed /ultrareview on macOS and Linux to require git 2.31 or newer to upload a local repository; checkouts made with --separate-git-dir are now refused instead of being uploaded with an older method
      • Changed /ultrareview uploads on macOS and Linux to send a partial clone as a working-tree snapshot on git 2.31 or newer, instead of falling back or refusing when git's version looked too old
      • Changed /ultrareview uploads on macOS and Linux to refuse, instead of fetching, a partial clone missing some of its working tree's files on older git versions; a clone made without --filter uploads
      • Changed Bedrock, Vertex and Mantle start-up model checks to identify themselves as Claude Code, like other Claude Code requests
      • Changed claude mcp get to hide the command, arguments, and environment values of stdio MCP servers provided by plugins; variable names are still shown
      • Changed /claude-api so it can no longer be run from Remote Control clients
      • Changed /config chrome=true to direct you to the /config panel instead of enabling Claude in Chrome by default; /config chrome=false still turns it off when it was on
      • Changed sandbox settings so project settings cannot widen or turn off an admin-required sandbox, replace the proxy behind a managed deny list, extend a strict allowlist, or reopen managed read-denies
      • [VSCode] Added a note under a restored tab's last message when a window reload interrupted it and no reply will follow
      • [VSCode] Added a plugin options form to Manage plugins: installing a plugin that has options asks for the unset ones, and a gear on its row changes them later
      • [VSCode] Added an on-demand diagnostics tool so Claude in the panel can read the Problems panel's current errors and warnings at any time, not only right after it edits a file
      • [VSCode] Fixed pressing Enter after typing a slash command running an unrelated menu item picked by fuzzy match, or doing nothing
      • [VSCode] Fixed an open agent transcript losing the agent's newer messages during a long session
      • [VSCode] Fixed a message that quotes a Claude Code or IDE tag losing the rest of its text in the chat
      • [VSCode] Fixed a message sent while Claude was working disappearing from the conversation after the session was reopened
      • [VSCode] Fixed the session list's Web tab showing the previous account's sessions after an account switch
      • [VSCode] Fixed restored tabs re-running an interrupted turn when VS Code was started with CLAUDE_CODE_RESUME_INTERRUPTED_TURN set, even with Continue After Reload off
      • [VSCode] Fixed Escape stopping the running turn instead of closing the command menu after clicking one of its rows
      • [VSCode] Fixed opening Past conversations replacing a live conversation with its saved copy
      • [VSCode] Fixed a Claude tab reloaded after an extension restart staying blank instead of saying how to recover
      • [VSCode] Fixed opening a conversation that is already open in another window or app starting a second copy of it without warning; it now asks first
      • [VSCode] Fixed a hook's reason for blocking or stopping a prompt disappearing after a window reload
      • [VSCode] Fixed tabs stuck on a conversation that can't be resumed: the error now says so and offers to start a new conversation
      • [VSCode] Fixed every file Read, Write and Edit stalling for ten minutes and then being skipped when the editor stops responding to the extension's automatic save before the tool runs
      • [VSCode] Fixed the agent map labeling a sub-agent with the session's model instead of the model it actually ran on (e.g. under CLAUDE_CODE_SUBAGENT_MODEL_FORCE or an agent's own model)
      • [VSCode] Fixed uninstalling a plugin from the Manage plugins dialog, which removed the wrong installation or failed for a plugin installed for the project
      • [VSCode] Fixed sign-in staying on the authorization-code step after going back and choosing the same sign-in method again
      • [VSCode] Fixed the chat panel stalling when a long session trims its oldest rows
      • [VSCode] Fixed the conversation disappearing from a session when many agents run
      • [VSCode] Fixed the editor tab keeping an old name after a session was renamed with /rename, by a SessionStart hook, or on claude.ai
      • [VSCode] Fixed the agent map's transcript view leaving out messages sent to a running agent
      • [VSCode] Improved the Manage plugins dialog: a failed plugin action now opens a popup that explains it and, where there is one, offers the fix
      • [VSCode] Changed the Manage plugins dialog to ask before removing a marketplace or turning off a plugin that your project's shared .claude/settings.json turns on
      • [Cloud sessions] Fixed Run now on a routine showing internal error text when the run is refused before it starts; it now shows the same explanation as the routine's failure notification
      • [Cloud sessions] Changed MCP_DISCOVERY_CACHE=1, when set in your cloud environment's variables rather than a settings file, to reuse your connectors' tool lists after a session restart; other MCP servers are no longer cached and connect at startup
      • [Claude Tag] Added direct messages with Claude for members on an Enterprise plan Standard or Usage-Based Chat seat who also have Cowork; a seat that includes Claude Code is no longer required
      • [Claude Tag] Fixed the Default model setting in admin settings and a channel's Configure page offering models your organization can't use, which made saves or new sessions fail
      • [Claude Tag] Fixed the note under Claude's Slack messages saying it answered on a fallback model, and why, disappearing when Claude later edited that message
      • [Code Review] Fixed the organization menu in Code Review's "Add a repository" dialog showing only a few of your GitHub organizations; it now loads more as you scroll
      • [Code Review] Improved the Code Review check run to say when your repository's REVIEW.md wasn't applied, for example on a very large pull request or when REVIEW.md is a symbolic link
    4. 🔗 earendil-works/pi v0.99.1 release

      New Features

      • GPT-6.1 Sol — Available on OpenAI, Azure OpenAI, and OpenAI Codex, and now the default OpenAI Codex model. See Select a model.

      Added

      • Added GPT-6.1 Sol (gpt-6.1-sol) to the OpenAI, Azure OpenAI Responses, and OpenAI Codex providers.

      Changed

      • Changed the default OpenAI Codex model to GPT-6.1 Sol (gpt-6.1-sol).

      Fixed

      • Fixed /login with OpenAI failing in the bundled release with a missing openai-chatgpt.js module error.
    5. 🔗 Evan Schwartz Please add prompt caching to Jev-style models rss

      If you're building a Jev-style "System One" model, please add prompt caching or reusable question sets to your API 🙏. This would make batch use cases even more efficient, so you could amortize the cost of many questions asked over the same input. (This was also proposed in typesafe-ai/typesafe-sdk-js#10.)

      TL;DR: after a week of tweaking my Jev calls, my questions are ~88% of the input tokens. I'm asking 54 questions of ~1.1 million documents per month. Jev makes certain types of classification tasks easy and cheap, but prompt caching would make batch workflows even more cost effective. For me, the total dollar amount is still reasonable (less than $150 per month), but I'm sure others will hammer these APIs even harder.

      Context

      I work on Scour. It’s a personalized content feed where you describe topics you’re interested in and it finds articles and blog posts related to them.

      For some time, I’ve wanted to ask slightly fuzzy questions of each piece of content Scour pulls in. What level of experience does this assume? Will this still be worth reading after a week, or is it news that goes out of date quicker? However, running the >1 million pieces of content Scour ingests each month through even the cheapest LLM would be prohibitively expensive for a bootstrapped project. System One models make this kind of use case feasible.

      TrAiN yOUr OwN cLaSsIfiEr

      Yes. I could. I might. But having a kind of general purpose classifier that I can ask a barrage of questions and add more questions to over time without retraining is very interesting.

      Tuning the question set

      I’ve spent a good chunk of the last week running experiments with Jev, tuning questions, comparing Jev's answers to panels of LLM judges and spot checking the results. The questions help me weed out junk that was hard to spot deterministically, figure out the expertise assumed by posts (to hide beginner content from advanced readers), and more.

      Following TypeSafe's advice about atomic questions and sending all questions in one request, v28 of my question set includes 54 questions: 50 nouls, 3 choices, and 1 score. I also reworded most of my questions to remove as much of the criteria description as I could without losing too much accuracy on the results.

      This set of questions is approximately 2,150 tokens, including Jev’s fixed overhead that seems to be around 260 tokens (judging from the token count for a trivially short question). When I send this set of questions in with the post’s title, URL, and summary or first snippet of text, my questions represent ~88% of my input tokens. The questions are sent verbatim with every request.

      Note that I don't send multiple posts in the same request because of the warning that "Accuracy falls as the state grows with content unrelated to the decision."

      Prompt caching or "Register-once questions"

      To TypeSafe and all the other labs that follow with these types of models, please add prompt caching or the ability to register sets of questions and criteria and reuse them.

      I don't know enough about how these models are served and whether there are caching efficiencies to be gained on the backend, but this would help increase Jev's effective "intelligence-per-dollar".

      With cheaper repeated questions, I'd put back some of the criteria I cut down, and probably add even more questions.

      As another alternative, an API that takes multiple questions and multiple inputs and gives you the results for each question evaluated for each input would give some of these savings without taking the accuracy hit from packing multiple inputs into one request.

      Batch mode?

      Async batch mode probably makes sense for super high-volume, fully offline use cases. It's less relevant for me because I want newly ingested content to be available soon after Scour reads it. But others will probably find plenty of use for it.

      Thanks

      Keep up the great work! This is a very exciting development.

    6. 🔗 r/LocalLLaMA Anthropic just dropped the greatest advertisement for GLM ever. rss

      Anthropic just dropped the greatest advertisement for GLM ever. | Like.. yea bro, I knew GLM was cool. Now everyone does. submitted by /u/BannedGoNext
      [link] [comments]
      ---|---

    7. 🔗 earendil-works/pi v0.99.0 release

      New Features

      • Codemode and MCP — Connect MCP servers and let models run JavaScript that calls tools in parallel. See MCP Servers and Enable codemode.
      • System theme — Pi's colors now come from your terminal's own palette by default. See Use your terminal's colors.
      • Sign in with ChatGPT — Use a ChatGPT subscription with the OpenAI provider through /login openai. See Authenticate interactively.
      • Virtual models — Extensions can route each request to a different physical model. See Virtual Models.
      • Classifier models — Run Jev classifiers from codemode scripts, or use any llama.cpp model as a classifier. See How codemode works and Classification.

      Added

      • Added codemode, tool search, and MCP support as built-in extensions. The codemode tool runs model-written JavaScript in a QuickJS sandbox that calls pi's tools; enable it with defaultTools or --tools and configure it with codemode.mode and codemode.inlineBudget. tool_search finds tools that are not declared to the model and declares them. MCP servers over stdio or streamable HTTP, with OAuth, come from mcp.json (global, or per project once trusted) or pi.registerMcpServer() and are managed with /mcp and pi mcp add|remove|list|login|logout. See MCP Servers and Enable codemode (#10040).
      • Added extension tool APIs for orchestrating tools: exposure (direct, model-only, codemode, deferred, or hidden), namespace, annotations, outputSchema with structuredContent, isError results, prepareLoadout(), and ctx.executeTool() for nested tool calls, which emit events with parentToolCallId and are recorded as bounded nestedCalls on the calling tool's result. See Tool exposure.
      • Added a warning when an extension that registers the same tool, command, or flag replaces a built-in extension (#10174 by @cristinaponcela).
      • Added experimental virtual models: extensions register them with pi.registerVirtualModel() and pick a physical model and thinking level for each request. The footer shows the routed model, /session lists cost per physical model, and examples/extensions/jev-router.ts routes with the Jev classifier. See Virtual Models.
      • Added Sign in with ChatGPT to the OpenAI provider in /login, which uses a ChatGPT subscription with the OpenAI API. Pi stores a stable deviceId in the global settings for this login and omits it from bug reports.
      • Added the system theme, now the default, which derives pi's colors from the terminal's reported foreground, background, and ANSI palette and rebuilds them when the terminal switches between light and dark. See Use your terminal's colors.
      • Added #rgb, oklch(), and okhsl() colors and an optional appearance field to theme files, and theme.style(), theme.colors, and theme.appearance for extensions. See Themes and TUI.
      • Added a classifier model for every llama.cpp chat model, answered from next-token label probabilities. See Classification.
      • Added inherited Jev classifier models on OpenRouter, Cloudflare Workers AI, Vercel AI Gateway, and OpenCode Zen.
      • Added the fullscreenWheelScrollLines setting and /settings entry for fullscreen mouse-wheel scrolling. The default "auto" accelerates fast wheel spins outside local macOS terminals (#9758).
      • Added per-input disposition to successful RPC prompt, steer, and follow_up responses, AgentSession.steer()/followUp(), and RpcClient.prompt()/steer()/followUp(); RpcClient.prompt() also accepts streamingBehavior (#9098, #9803).
      • Added image generation to ModelRuntime: generateImages() with runtime-resolved auth (stored credentials, OAuth, runtime API keys, models.json headers), plus getModelsOfType(), getModelOfType(), getAvailableOfType(), getAllModels(), and getAllAvailable(). OpenRouter image models are listed under the openrouter provider and share its credential; an upstream ID can have separate chat and image entries. models.json providers and extension registrations without a model list keep built-in image generation. Extension model lists can include discriminated chat, image, and classifier entries with operation implementations; when supplied, they replace the provider catalog across every operation. Chat-facing reads (getModels(), getAvailableSnapshot(), the model picker) are unchanged.
      • Added classifier support to ModelRuntime, including classify(), classifier model accessors, runtime-resolved authentication, and the built-in TypeSafe jev-latest model.
      • Added types=chat,image,classifier to pi.dev model catalog requests so remote refreshes overlay every supported model type; entries of unknown model types are ignored.
      • Added the provider_stream_event extension event for observing parsed provider events before normalization, with an opt-in /debug-provider example viewer (#9784, #9901 by @davidbrai).
      • Added a show/hide toggle (H) in HTML exports for custom messages marked display: false. Messages remain hidden by default and can also be revealed from the sidebar (#8896, #10020 by @rwachtler).
      • Added inherited Claude Sonnet 5.5 support for Anthropic with adaptive thinking and a 1M context window.
      • Added a Built-in section in pi config to disable the built-in mcp, llama.cpp, codemode, and tool-search extensions globally or per project, stored as -builtin:<name> in the extensions setting. SDK inline extensions opt in with builtin: true.
      • Added +name and -name entries to the defaultTools setting to add or remove tools without repeating the defaults, for example "defaultTools": ["+codemode"]. Project entries of this form apply on top of the user setting. Documented how to enable codemode without MCP and how to use classifier models such as Jev from codemode scripts.
      • Added the token usage and cost of codemode models.classify() calls to the codemode tool result, so they count toward the session cost; the codemode result shows each call's cost.

      Changed

      • Switched the build from the TypeScript native preview to TypeScript 7.0 with an ES2024 target, and replaced tsx with Node's built-in type stripping for running from source (#9965).
      • Removed the [Themes] section from the startup banner. Custom themes remain available in /settings, and theme conflicts are still reported.
      • Changed the startup header to show the pi logo with the version instead of the app name.
      • Changed the built-in dark and light themes to the revised pi colors, written in OKHSL.
      • Changed light/dark terminal detection to use the reported background color first, then the terminal's light/dark report, then COLORFGBG. The first-time setup no longer shows the detected appearance.
      • Renamed the inherited OpenAI Codex provider to "OpenAI Codex (legacy)"; Sign in with ChatGPT on the OpenAI provider supersedes it.
      • Changed inherited terminal detection to treat TERM=*-direct as truecolor.
      • Built-in extensions and tools are named builtin:<name> (for example builtin:mcp and builtin:read) in errors, diagnostics, RPC source info, and bug reports, instead of <inline:name> and <builtin:name>. Their slash commands no longer carry a [t] autocomplete tag.
      • --no-extensions also disables the built-in extensions, including the llama.cpp provider. Load one explicitly with -e builtin:<name>, for example pi -ne -e builtin:mcp.
      • Tool calls without a custom call renderer, including direct MCP tool calls, now show their arguments: as key=value pairs on the title line when collapsed and one key: value line per argument when expanded. MCP calls are titled server/tool and their results collapse to 5 lines.
      • bash and powershell structured results, which codemode scripts receive, now hold up to 1 MiB of output instead of the model-facing 2000 lines or 50KB, and add truncated and full_output_path. Longer output keeps its first and last 512 KiB. Empty output is "" instead of (no output).

      Fixed

      • Fixed X11 clipboard text being misidentified as an image when the clipboard owner accepts unadvertised image targets (#9786).
      • Prevented managed git packages from automatically installing Pi peer dependencies and added warnings for extension packages that list host-provided modules in dependencies (#9863).
      • Fixed pinned git extensions loaded with -e continuing to use the first downloaded commit after the ref changes (#9982).
      • Fixed RpcClient skipping the next event listener when a listener unsubscribes while handling an event, which could make waitForIdle() time out after collectEvents() (#9990).
      • Fixed full-file read calls rendering as :1 when models send null for omitted offset and limit (#9996).
      • Fixed new sessions being lost when pi exits before the first assistant response. The session file is now created when the first user message is sent (#10000).
      • Fixed unloaded llama.cpp autoload presets overwriting a cached runtime context window with the GGUF training context (#10077, #10158 by @cristinaponcela).
      • Fixed custom themes ignoring terminal.trueColor and other terminal capability overrides and rendering with 256 colors (#9973, #10039 by @christianklotz).
      • Fixed pasting files copied in Finder inserting the file icon image instead of the file paths; paths are quoted in bash mode (#9999, #10136 by @christianklotz).
      • Fixed the startup header, loaded resources, and chat notices keeping their old colors after a theme change.
      • Fixed the Fireworks default model pointing at the removed Kimi K2.6 model; it now defaults to Kimi K3.
      • Fixed the OpenCode Go default model pointing at the removed Kimi K2.6 model; it now defaults to Kimi K3.
      • Fixed the Together default model pointing at the removed Kimi K2.6 model; it now defaults to Kimi K3.
      • Reduced CPU use while streaming in long sessions and when previewing themes: the footer caches session usage totals, collapsed bash results cache their preview, and sanitizeBinaryOutput() no longer splits output into per-character arrays.
      • Fixed the usage of tools called through ctx.executeTool(), for example from codemode scripts, being dropped from the session cost; it is now added to the calling tool's result usage.
      • Fixed inherited /skill autocomplete appearing empty when loaded skill names did not contain the letters in skill (#9944).
      • Fixed inherited path and @ autocomplete not working after opening wrappers such as (, [, {, <, or a backtick.
      • Fixed inherited image stretching in terminals that use the Kitty graphics protocol (#8938, #9957 by @rwachtler).
      • Fixed inherited shell cursor staying hidden after exit when an extension closed an overlay during shutdown (#10026).
      • Fixed inherited keyboard input being lost after a mouse click in a /settings submenu closed it.
      • Fixed inherited 1-hour Anthropic cache writes through Vercel AI Gateway being priced at the 5-minute rate (#9210).
      • Fixed inherited model-level samplingParams being dropped by direct stream()/complete() calls on OpenAI-compatible APIs (#9506).
      • Fixed inherited Mistral GLM requests failing with "Expected at most one leading ThinkChunk" after empty content deltas (#9674).
      • Fixed inherited OpenAI Fast mode requests being priced at the standard rate (#10034).
      • Fixed inherited Mistral reasoning models ignoring the requested thinking level (#9678).
      • Fixed inherited OpenCode Zen and OpenCode Go qwen3.8-flash thinking being replayed as plain text on later turns (#10047).
      • Fixed inherited OpenAI Responses streams from servers that omit output_index, such as llama.cpp, running mixed-up tool calls; such streams now end with an error (#9974).
      • Fixed inherited Anthropic and OpenAI Codex browser sign-in waiting indefinitely after the provider redirected with an authorization error, and Anthropic sign-in failing when its callback port is in use.
      • Fixed inherited GitHub Copilot Claude Opus 5.5 offering unsupported thinking levels when upstream model metadata is incomplete.
    8. 🔗 r/LocalLLaMA Deepseek Harness app is out now!!! rss

      Deepseek Harness app is out now!!! | Downloading now. submitted by /u/politefella0
      [link] [comments]
      ---|---

    9. 🔗 r/LocalLLaMA GLM-5.3 and the Spread of Advanced Cyber Capabilities \ Anthropic rss

      GLM-5.3 and the Spread of Advanced Cyber Capabilities \ Anthropic | submitted by /u/brown2green
      [link] [comments]
      ---|---

    10. 🔗 HexRaysSA/plugin-repository commits sync repo: +1 plugin, +2 releases rss
      sync repo: +1 plugin, +2 releases
      
      ## New plugins
      - [mcrit-ida](https://github.com/familiary/mcrit-plugin) (2.0.0)
      
      ## New releases
      - [ida-mcp](https://github.com/hexrayssa/ida-mcp): 20260929.0.1
      
    11. 🔗 Simon Willison OpenAI DevDay 2026 live blog rss

      I'm at OpenAI DevDay today, in Fort Mason, San Francisco. Same as last year I'll be live blogging the keynote and some other notes during the day.

      OpenAI gave me a free ticket and a seat in the "creator" area for the keynote.

      You are only seeing the long-form articles from my blog. Subscribe to /atom/everything/ to get all of my posts, or take a look at my other subscription options.

    12. 🔗 r/LocalLLaMA AMD's new 256 core EPYC has 16-channel DDR5-12800, 91% memory bandwidth of an RTX 5090 rss

      AMD's new 256 core EPYC has 16-channel DDR5-12800, 91% memory bandwidth of an RTX 5090 | Here are some inspirational quotes you can put into the comments:

      • God is dead and we killed him
      • I am become death
      • All this for 1.5 tok/s?
      • Sir this is LocalLLaMA not RichPeopleofLocalLLaMA
      • Sweet! A 2TB DDR5-12800 RDIMM kit is going to cost only 2 kidneys and a small micronation's GDP

      submitted by /u/Dany0
      [link] [comments]
      ---|---

    13. 🔗 roboflow/supervision supervision-0.30.6 release

      0.30.6: Pillow-conversion, dataset-export, and line-crossing correctness

      fixes

      supervision 0.30.6 is a patch release closing 12 library correctness bugs and landing 2 behavior refinements across datasets, annotators, key points, line- crossing counting, and the Pillow-based fallback backend, plus a new docs guide and a notebook dead-link fix. No new public API, no breaking changes. The broadest-reach fix rewrites pillow_to_cv2, the entry point every annotator except RichLabelAnnotator and several image helpers use to convert a Pillow image: it used to pass raw mode bytes straight through, so a 1-bit mask, a 16-bit depth map, an LA/PA image, or a CMYK JPEG all drew wrong or crashed. Two fixes close silent-wrong-data bugs rather than crashes: LineZone.trigger let unconfirmed tracks inflate crossing counts, and DetectionDataset.as_yolo/.as_pascal_voc silently dropped box-only objects from a mixed polygon/box COCO dataset. DetectionsSmoother no longer crashes when two tracked objects first seen on different frames carry different metadata (e.g. RF-DETR's source_image); its own docstring pipeline was broken. Every library fix ships a regression test.

      ✨ Spotlights / highlights

      sv.pillow_to_cv2 now converts every Pillow mode the way cv2.imread

      would

      Raw mode bytes used to pass straight through: a 1-bit mask came back as 0/1 instead of 0/255, a 16-bit depth map wrapped modulo 256 and drew as noise, an LA/PA image crashed in cvtColor on its 2-channel array, and a CMYK JPEG drew its cyan/magenta/yellow ink as red/green/blue. This function is the entry point for every annotator except RichLabelAnnotator (which stays on the Pillow path) plus crop_image, resize_image, letterbox_image, scale_image, tint_image, grayscale_image, and plot_image whenever handed a Pillow image, so the bug reached the whole drawing surface, not just one call site. A single-channel grayscale image also drew from a read-only buffer view before this fix. Annotating into it silently failed; it now hands back a writable copy.

      from PIL import Image
      
      depth_map = Image.open("depth.tif")  # mode "I;16"
      scene = sv.pillow_to_cv2(depth_map)
      # clipped to the 16-bit range and keeps its high byte, instead of wrapping mod 256
      

      RGB, RGBA, grayscale, and palette images are unchanged.

      sv.LineZone.trigger no longer lets unconfirmed tracks inflate counts

      Detections with a negative tracker_id (how trackers like ByteTrackTracker mark an unconfirmed track) all keyed under one shared id. Two distinct unconfirmed objects crossing on opposite sides in the same frame read as one track oscillating, silently inflating in_count/out_count with crossings no confirmed track ever made. Confirmed tracks (tracker_id >= 0) are counted exactly as before.

      sv.DetectionDataset.as_yolo / .as_pascal_voc stop dropping box-only

      objects

      from_coco gives an all-zero mask to any annotation without its own segmentation when a sibling annotation has one, and to every annotation when force_masks=True. The exporters only ever wrote the polygon traced from a mask, so a mixed polygon/box COCO dataset silently lost its box-only objects on export, and a box-only COCO dataset loaded with force_masks=True wrote empty YOLO label files or object-less Pascal VOC files. A detection with an empty or contour-less mask now exports as its bounding box instead.

      sv.DetectionsSmoother no longer crashes on a second tracked object

      Each smoothed track was built on the oldest frame in its window, so two objects first seen on different frames carried different frames' metadata (exactly what the RF-DETR / inference connectors attach), and Detections.merge rejected the mismatch. class_id and data (e.g. class_name) now follow the current frame instead of lagging up to length - 1 frames behind a class change; xyxy, confidence, and oriented-box corners are still averaged as before.

      Dataset splits reject an out-of-range ratio instead of silently mis-

      splitting

      sv.DetectionDataset.split, sv.ClassificationDataset.split (both take split_ratio), and the internal train_test_split helper they call (takes train_ratio) never validated the ratio. A finite out-of-range value looked like a successful split: for 10 images, a ratio of -0.2 returned 8/2 via negative slicing, and 1.2 (or an accidental percentage like 80) returned 10/0 with no held-out data, no error either way.

      train, test = dataset.split(split_ratio=80)  # meant 80%
      # now raises ValueError naming the problem, instead of silently returning
      # every image as train and none as test
      

      0 and 1 keep their existing meanings. Not an API break: no signature changed, only previously-silent input now raises.

      🔄 Migration guide

      No breaking changes in this release.

      No deprecations or removals landed in 0.30.6 either. All scheduled remove_in markers in the codebase target 0.31.0 or 0.32.0 and are untouched by this patch release.

      Two entries change output values rather than only fixing a crash or a silently-wrong count; not API breaks, but worth checking if your code depends on the old values:

      • sv.KeyPoints.from_ultralytics / .from_inference / .from_detectron2: as_detections().confidence now carries the model's own detection score, not the mean of the per-keypoint confidences.
      • sv.DetectionsSmoother: class_id and data fields on a smoothed track now follow the current frame instead of lagging behind a class change.

      📝 Notable changes

      🔧 Fixed

      Annotators / key points

      • sv.VertexEllipseHaloAnnotator now draws the whole halo of a key point whose covariance ellipse is not horizontal. A vertical or diagonal ellipse used to be clipped to a thin band because the fade box was sized as if the major axis always ran along the image x axis. Horizontal ellipses are drawn exactly as before. (#2634)
      • sv.KeyPoints.from_ultralytics, .from_inference and .from_detectron2 now keep each object's detection score as detection_confidence instead of dropping it, which previously made sv.KeyPoints.with_nms raise ValueError and as_detections() report the mean keypoint confidence instead of the model's own score. (#2633)
      • sv.LabelAnnotator now sizes the label background correctly for a label containing a blank line. Text drew each blank line at full height while the background measured it as zero, pushing the last line outside the box. Labels without blank lines are unchanged. (#2632)
      • sv.IconAnnotator now draws a palette icon that carries its own alpha channel (Pillow PA mode, storable in TIFF) instead of raising ValueError. Palettes without alpha, and every read with OpenCV installed, are unchanged. (#2624)

      Datasets

      • sv.DetectionDataset.as_yolo / .as_pascal_voc now write a detection whose mask is empty or has no valid contour as its bounding box instead of silently leaving it out of the label file. (#2631)
      • sv.DetectionDataset.from_yolo now loads a label row carrying a trailing confidence or tracker id (written by Ultralytics save_txt with save_conf=True or tracking on) instead of raising ValueError, for both box and segmentation rows. The extra token is ignored; box and polygon geometry are unchanged. An odd-length polygon row with no extra field now loses its last value rather than raising, since coordinate parity is the only signal separating the two cases. (#2619, #2626, #2635)

      Tracking

      • sv.DetectionsSmoother no longer raises ValueError: Conflicting metadata when a second tracked object, first seen on a different frame, enters the smoothing window. (#2628)
      • sv.LineZone.trigger now ignores detections with a negative tracker_id (the value trackers such as ByteTrackTracker report for an unconfirmed track) instead of letting them silently inflate in_count/out_count. (#2623)

      Detection utils

      • sv.polygon_to_mask now accepts a list/tuple/array-like of [x, y] vertices (not just a NumPy array) and returns an all-zero mask for an empty or under-3-vertex polygon instead of crashing inside OpenCV/NumPy with an opaque error. A malformed polygon now raises ValueError naming the problem. (#2622)
      • sv.Detections.from_sam3 now keeps SAM 3 PVS contour fragments with fewer than 3 vertices instead of dropping them. Single points and 2-point edges are rasterized directly into the mask and bounding box. (#2625)

      Image / IO

      • sv.pillow_to_cv2 now converts every Pillow mode to the 8-bit array cv2.imread would produce, reaching every annotator plus crop_image, resize_image, letterbox_image, scale_image, tint_image, grayscale_image, and plot_image. (#2614)
      • sv.CSVSink now writes UTF-8 on every platform, preserving non-English detection labels and custom fields on Windows. (#2615)

      Docs

      • Dead documentation links fixed across the published notebooks (quickstart.ipynb, annotate-video-with-detections.ipynb, underestand-visitors-with-yolo-world.ipynb). (#2630)

      🌱 Changed

      • sv.DetectionDataset.split, sv.ClassificationDataset.split, and the internal train_test_split helper they call now reject a ratio outside [0, 1] (NaN and ±inf included) with a ValueError naming the problem, instead of silently mis-splitting or dropping the held-out set. (#2611)
      • sv.tint_image and sv.grayscale_image now accept a single-channel (H, W) array or grayscale Image, as sv.letterbox_image already did. (#2614)
      • New docs guide: "Evaluate targets from a PyTorch DataLoader" in docs/metrics/mean_average_precision.md, cross-linked from docs/how_to/benchmark_a_model.md. Detections.from_transformers gained a docstring note that it expects post-processed predictions, not raw DataLoader targets. (#2620)

      🏆 Contributors

      • Mohammad Hijjawi (@MohammadHijjawi97, LinkedIn) — the DetectionsSmoother conflicting-metadata crash, the KeyPoints detection-score fix, the YOLO/Pascal VOC empty-mask export fix, the VertexEllipseHaloAnnotator rotated-ellipse fix, the LabelAnnotator blank-line background fix, and the polygon_to_mask crash fix
      • Raghav Rathi (@raghav-rathi, LinkedIn) — the pillow_to_cv2 rewrite for every Pillow mode
      • Zachary Alexander (@zachsplat) — the LineZone unconfirmed-track fix
      • NIKHIL (@Nikhi00718) — the YOLO segmentation-label trailing-value fix
      • Manoj Kumar Thapa (@iammanoj807, LinkedIn) — the palette-with-alpha IconAnnotator fix
      • kevin (@kevin9327) — the YOLO box-label trailing-value fix
      • Andrew Barnes (@Bortlesboat, LinkedIn) — the CSV Unicode export fix
      • 4rch1e (@archie0732) — the dataset split-ratio validation
      • Salman Ansari (@salmanwnl44, LinkedIn) — the PyTorch DataLoader-targets docs guide
      • Pratik Gandhi (@pratikgx, LinkedIn) — dead-link fixes in the published notebooks
      • Jirka Borovec (@Borda, LinkedIn) — the from_sam3 degenerate-fragment fix

      Full changelog : 0.30.5...0.30.6

    14. 🔗 r/LocalLLaMA Looks like the era of subsidised compute is coming to an end. The old ChatGPT Pro $200 20x plan will be halved. The new $500 plan will have similar limits as the (old) $200 plan. rss
    15. 🔗 smol-machines/smolvm smolvm v1.20.2 release

      What's Changed

      • Resize CPUs, RAM and disks of a running machine in one command by @BinSquare in #1466
      • Let a stopped machine's egress policy be replaced by @BinSquare in #1449
      • Give a machine that is still starting time before the ephemeral sweep reaps it by @BinSquare in #1470
      • Bump the Nix package to 1.20.1 by @BinSquare in #1469
      • End a foreground packed run's VM with its CLI, and watch its mounts from the CLI by @BinSquare in #1467
      • Bump the workspace to 1.20.2 by @BinSquare in #1468

      Full Changelog : v1.20.1...v1.20.2

    16. 🔗 smol-machines/smolvm smolvm v1.20.1 release

      What's Changed

      Full Changelog : v1.20.0...v1.20.1

    17. 🔗 backnotprop/plannotator v0.27.22 release

      Follow @plannotator on X for updates

      Missed recent releases? Release | Highlights
      ---|---
      v0.27.21 | Remote and phone sessions load several times faster, real Request changes on GitHub, model pickers show real names, OpenCode fixes
      v0.27.20 | Mistral Vibe support, annotate gets the full Options menu and Settings, jj Commits panel, long lines wrap in plan code blocks
      v0.27.19 | Before/After image previews in code review, file comments as GitHub file threads, forge-correct #123 links, /plannotator-last finds the right session
      v0.27.18 | Model pickers from your installed Claude and Codex (Opus 5.5, Fable 5.1, GPT-6), unsent PR review comments survive new pushes
      v0.27.17 | Diagram files open in the diagram viewer, OpenCode switches model with agent, idle review stops polling the git remote, Tree is the default review view
      v0.27.16 | Themed diagrams on Mermaid 12, comment on any node or edge, patch-file review, embedded HTML documents render
      v0.27.15 | Plannotator TUI and Herdr Annotate announcement, element context on pinpoints, HTML links open as linked documents, All files panel, Classic diff default
      v0.27.14 | Pi plan progress survives compaction, Codex threads across rollout files, WSL browser setting, Mod+E edit mode
      v0.27.13 | Open a review on a specific base (--base, --diff-type), symlink containment on /api/doc, CI flake fix, Amp decision relay
      v0.27.12 | Unified decision control, token hover cards, local-vs-remote diff, approval notes
      v0.27.11 | OpenCode server leak fix, durable local feedback archive, unknown-subcommand fix
      v0.27.10 | Auto-viewed files on scroll, annotation undo/redo, OpenCode 2 slash commands restored, npm 12 agent terminal fix

      What's New in v0.27.22 A fix release. Plans can open the documents they link to, Claude review jobs are locked down more tightly, Code Tour works on Linux machines where Claude Code's sandbox cannot start, and Pi no longer reviews an outdated plan. Eight pull requests, two of them from community contributors, answering four reports. Plans can open the documents they link to Claude Code saves plans in ~/.claude/plans/, outside your project. Plannotator only looked inside the project for linked files, so a plan that linked or [[notes]] to a file saved beside it showed a link that went nowhere. Plan review now also looks in the plan file's own folder, but only for files the plan actually links to, matched by exact name. Every plan you have ever made lives in that folder, so a plan can open its own notes without being able to reach any other plan. Links inside code blocks do not count, and plans kept inside the project still fall back to the project root as before. If a file beside the plan has the same name as a project file, the one beside the plan wins. (#1620, closing #1443, by @bendrucker) Claude review jobs are locked down

      Code review, Code Tour and Guided Review run Claude Code with a list of read- only commands they are allowed to use. That list had grown looser than its purpose: a few entries accepted options that could run other programs or write files, and the review prompt contains the diff under review, which is untrusted text. For most people, Claude Code's own sandbox is off, so that list was the only thing standing between a hostile diff and the machine.

      The list is now narrowed to the commands the prompts actually use, shared by all three job types, with the risky options explicitly refused. Jobs also no longer read a repository's own .claude/settings.json, so a pull request you are reviewing cannot add permissions or hooks to the job. Your personal Claude Code settings and any managed settings still apply. Jobs no longer load your MCP servers either; before, a Code Tour could wander into tools from your IDE.

      (#1628, #1629)

      Code Tour on Linux when Claude Code's sandbox cannot start

      If your Claude Code settings enable its sandbox and the machine lacks bubblewrap and socat, the sandbox fails to start and Claude refuses every command. Code Tour then ended with "Tour blocked: git access is unavailable" and no hint why.

      Now, when a Claude job has every ordinary git command refused, the Agents tab shows a warning explaining what happened. You can install those two packages, or turn Claude's sandbox off for Plannotator's jobs only with PLANNOTATOR_CLAUDE_SANDBOX=0 or { "claudeSandbox": false } in ~/.plannotator/config.json. Plannotator never turns the sandbox off on its own, since that would quietly override a choice you made.

      (#1628, closing #1627, reported by @Infeligo)

      Pi reviews the plan you just wrote

      Pi runs all the tool calls in one assistant message at the same time. When the agent edited the plan and submitted it in the same message, the submit could read the file before the edit landed, so the review opened on the previous version and version history recorded it. The submit tool now tells Pi to run its batch in order, so the edit always finishes first. Every Pi version the extension supports honors this.

      (#1623, closing #1622, reported by @bugs-wkettlitz)

      Additional Changes

      • Claude reviews no longer fail after producing findings. Recent Claude Code versions can end one run with several result messages when the model uses background subagents, and the last one is often empty. Plannotator read only the last message, so about one code review in three was marked failed even though Claude had returned findings. Review, Code Tour and Guided Review now read the newest message that carries results (#1630).
      • VS Code tabs follow a changed data folder. Plannotator looked up VS Code's connection file once at startup, so setting PLANNOTATOR_DATA_DIR later made opening a review in VS Code quietly fall back to the browser. It now looks the file up each time (#1621, closing #1502, by @gwynnnplaine).
      • Tests stay out of your data folder. Running the test suite from a package directory wrote drafts, history and saved plans into the real ~/.plannotator, and a personal config could make tests fail. Every package now uses the same temporary data folder the root run does (#1624, #1626).

      Install / Update

      macOS / Linux:

      curl -fsSL https://plannotator.ai/install.sh | bash
      

      Windows:

      irm https://plannotator.ai/install.ps1 | iex
      

      Claude Code Plugin: Run /plugin in Claude Code, find plannotator , and click "Update now".

      Pi: Update @plannotator/pi-extension to 0.27.22 and restart Pi.

      OpenCode: Clear cache and restart:

      rm -rf ~/.bun/install/cache/@plannotator
      

      What's Changed

      • fix(server): resolve the VS Code IPC registry path per call by @gwynnnplaine in #1621
      • fix(plan): open docs linked beside the plan file by @bendrucker in #1620
      • fix(pi): run plannotator_submit_plan sequentially so it never reads a stale plan by @backnotprop in #1623
      • test(pi): sandbox review-server tests from the user's config by @backnotprop in #1624
      • test: sandbox the data dir when tests run from a package directory by @backnotprop in #1626
      • fix(agents): hermetic Claude review jobs + sandbox opt-out for Linux by @backnotprop in #1628
      • security(agents): narrow Claude job command allowlists to what the prompts use by @backnotprop in #1629
      • fix(agents): read structured output from any Claude result event, not just the last by @backnotprop in #1630

      Contributors

      @bendrucker made plans able to open the notes and documents saved next to them, a gap every Claude Code user with multi-file plans ran into.

      @gwynnnplaine tracked down the last module that froze the data folder at startup and fixed it with tests that reproduce the bug.

      Community

      • @bugs-wkettlitz diagnosed the Pi race down to pi's parallel tool execution and proposed the exact fix in #1622
      • @Infeligo reported Code Tour failing on Linux, with the full tool log, in #1627

      Full Changelog : v0.27.21...v0.27.22

    18. 🔗 HexRaysSA/plugin-repository commits sync repo: -1 plugin, -4 releases rss
      sync repo: -1 plugin, -4 releases
      
      ## Removed plugins
      - mcrit-ida
      
    19. 🔗 smol-machines/smolvm smolvm v1.20.0 release

      What's Changed

      • Bump the Nix package to 1.19.3 by @BinSquare in #1441
      • Add opt-in exact and wildcard egress host patterns by @ankrgyl in #1438
      • Keep a paused machine's disks in place, and share backing copies across pauses by @LoganGrasby in #1442
      • Stream sparse RAM into stored checkpoints by @BinSquare in #1443
      • Quote --at '~N' in hints and docs so zsh doesn't expand it by @BinSquare in #1446
      • Stop decoding pause archives after resume inputs by @BinSquare in #1445
      • Remove duplicate reads and RAM copies from pause resume by @BinSquare in #1447
      • Parse checkpoint indexes from memory so saves and restores stop slowing down as history grows by @BinSquare in #1448
      • Layer a macOS branch over a restored disk with an absolute backing path so it can boot by @BinSquare in #1451
      • Bump libkrun to live resource resizing and branch pause fixes by @BinSquare in #1450
      • Create machines from an unchanged private pack without re-reading it by @BinSquare in #1455
      • Stop re-extracting VM-mode packs on every launch by @BinSquare in #1456
      • Let a non-root resume ignore the root-only restore tmpfs it cannot read by @BinSquare in #1452
      • Bump libkrun to the macOS branch pause fix by @BinSquare in #1460
      • Re-extract a pack entry whose image layers were removed after extraction by @BinSquare in #1459
      • Grow running machines without rebooting by @BinSquare in #1319
      • Boot the children of concurrent branches from one frozen source together by @BinSquare in #1457
      • Pause packed machines in place and stop their pause from crashing the VM by @BinSquare in #1461
      • Bump the workspace to 1.20.0 by @BinSquare in #1462

      Full Changelog : v1.19.3...v1.20.0

    20. 🔗 Filip Filmar Vreteno: a RISC-V core written in TxHDL rss

      Vreteno is a 32-bit RISC-V processor core designed in TxHDL. TxHDL is a hardware description language implemented as a native Rust library. The core implements the RV32IMC instruction set architecture: the base integer set, standard multiply/divide extensions, and compressed instructions. It supports traps, hardware interrupts, and system memory access via an AXI bus. Lockstep verification checks core state against a reference model on every cycle, while gate-level netlists match cycle-accurate Rust simulations. This article details the core microarchitecture, verification strategy, and FPGA implementation.

    21. 🔗 Filip Filmar wdbcvt: Bundling Waveform Test Corpora and Daily Nightly Releases rss

      wdbcvt now publishes two new release archives with every build: a documentation bundle and a waveform corpus. The waveform archive packages the test suite’s .wdb files, ground-truth .fst conversions, and HDL source files together. Nightly builds now run daily, check the repository forge for new commits, and publish all build artifacts to Forgejo and GitHub.

      Why a standalone waveform corpus matters

      Tool developers who study Vivado’s .wdb format need test cases. Vivado is a multi-gigabyte software package. Running simulations requires licenses and tool setups that many open-source developers do not maintain. The wdbcvt repository already contains dozens of simulated test cases across VHDL, Verilog, and SystemVerilog. Previously, these test cases stayed inside the repository tree alongside intermediate simulation files. Developers of external tools had to clone the full repository or run Vivado locally to obtain sample waveforms.

    22. 🔗 Armin Ronacher Deser: Rethinking Rust Serialization rss

      Serde is an amazing serialization library for Rust and it has been a huge reason why I felt productive with it for years. However already while at Sentry I got quite frustrated with some of the limitations with it but actually replacing Serde is tricky because of the might that it has in the ecosystem. Also because it's quite hard to actually do better without also making some potentially painful compromises.

      Here are three examples of Serde corner cases that show poor interactions of Serde features or unexpected limitations:

      A number that is a map

      An internally tagged enum, with serde_json's arbitrary_precision feature turned on:

      #[derive(Deserialize)]
      #[serde(tag = "type")]
      enum Shape {
          Circle { radius: f64 },
      }
      
      serde_json::from_str::<Shape>(r#"{"type": "Circle", "radius": 1.5}"#)
      // error: invalid type: map, expected f64
      

      Serde's data model has no place for arbitrary precision numbers, so serde_json uses in-band signalling with a map with a magic key. The enum has to buffer the fields until it has seen the tag, and the buffer does not know about the magic key. Because Cargo features are unified, it's enough for any crate in your dependency graph to turn the feature on.

      Flattening breaks integer keys

      #[derive(Deserialize)]
      struct Stats {
          scores: HashMap<u32, u32>,
      }
      
      #[derive(Deserialize)]
      struct Report {
          name: String,
          #[serde(flatten)]
          stats: Stats,
      }
      
      serde_json::from_str::<Report>(r#"{"name": "x", "scores": {"42": 23}}"#)
      // error: invalid type: string "42", expected u32 at line 1 column 35
      

      Stats on its own parses {"scores": {"42": 23}} just fine. JSON keys are always strings, and serde_json only turns them into integers if the type asks for one. However once flatten buffers the value, "42" is just a string. The error also points at the end of the document rather than at the key.

      Adapters do not compose

      fn from_hex<'de, D: Deserializer<'de>>(d: D) -> Result<u32, D::Error> { ... }
      
      #[derive(Deserialize)]
      struct Theme {
          #[serde(deserialize_with = "from_hex")]
          primary: u32,
          #[serde(deserialize_with = "from_hex")]
          accent: Option<u32>,
      }
      
      //error[E0308]: `?` operator has incompatible types
      //  |
      //  |     #[serde(deserialize_with = "from_hex")]
      //  |                                ^^^^^^^^^^ expected `Option<u32>`, found `u32`
      //  |
      //help: try wrapping the expression in `Some`
      //  |
      //  |     #[serde(deserialize_with = Some("from_hex"))]
      //  |                                +++++          +
      

      A function cannot be passed as a type parameter, so there is no way to apply from_hex to the inside of an Option, a Vec or a map. You write another function for every wrapper, and once you have from_opt_hex the field is no longer optional unless you also remember to add #[serde(default)].

      None of these are bugs that are easy to fix in Serde. They fall out of its design, and that design is protected by Serde's stability guarantees.

      Back in 2022 I started an experiment called Deser. It's a serialization library for Rust that takes the user experience of Serde and puts it on top of a completely different architecture inspired by miniserde. I never really finished it and it sat around for a few years. I picked it back up, and it has now reached a point where I think it's worth looking at. Even just to inspire others to see if they want to explore the space.

      The Name And Idea

      The name is Serde with its two halves swapped. Deser is Serde but the other way around. In Serde, a type drives the deserialization process: a Deserialize impl asks the deserializer for the kind of value it expects, the format calls back into a visitor. Every nested value is handled by recursion which makes Serde deserialization inherently grow the stack with each level of nesting.

      Deser on the other hand turns this around and the format tells the type of the next value and pushes events into a sink. When a sink hits the start of a nested value, it doesn't call into it but hands back a new sink to a driver, which keeps all state on the heap (in fact, in an arena). On the way out, emitters return their nested values instead of recursing into them.

      That also means that Deser cannot support formats like protobuf that are not self describing. They are in fact quite intentionally left out of the design entirely. Which is one way to say: if you want to "fix" Serde, you need to make some other compromises.

      Most of the reasons for Deser's ideas go back to Sentry Relay, which processes enormous amounts of untrusted JSON. Over the years when I was at Sentry we ran into the same set of problems again and again, and many of them are not really bugs in Serde but consequences of its design. Serde's stability guarantees mean that a lot of them cannot be fixed without breaking every format and every hand written implementation. Most of these problems come from three decisions:

      1. One set of traits for all formats. Serde serves both self describing formats (JSON, YAML, TOML, …) and formats where the reader has to know the type upfront (postcard, bincode, protobuf, …). That is incredibly useful, but it means that some features only work with some formats, and you find out at runtime. In case of Serde it also has some odd wrinkles where a derived struct quietly accepts an array in place of an object in JSON for instance.

      2. A fixed data model that loses information when buffering. Internally tagged enums, untagged enums and flatten need to buffer values before they know what to do with them. The buffer can't hold everything the format knew, errors lose their location and extensions to the ecosystem rely on in-band signalling to express things such as arbitrary precision numbers.

      3. Recursion on the call stack. Every level of nesting uses stack space. Formats protect against this with a recursion limit, but the moment you go through a code path that doesn't have one (writing, dynamic values), deeply nested data can take down your process. It also means that a deserialization cannot be paused while you wait for more input.

      Many of the corresponding Serde issues have been open for years, and I wrote about abusing Serde before. People have tried different angles on this over the years. Some went minimal and dropped most features to get fast compiles and no recursion. dtolnay's own miniserde is the best example of that, and deser's trait design was originally modelled after it. Other recent attempts went for runtime reflection, or for a new data model with a focus on binary formats.

      If you want to read up on all of the collected challenges with Serde's design, I maintain a lengthy list here.

      Dethroning Serde

      First of all I don't think it's likely that one can replace Serde. The orphan rule entrenches Serde incredibly well in the ecosystem. But some things are within the reach of a crate author's control. In case of Deser it's completeness.

      Deser today implements all important self describing formats from YAML, JSON, TOML, CBOR, JSON5 and the likes, but also XML and plist to really close the gap. XML in particular is something Serde has declined to support, and it shows (more on that below). At the very least format support should not be the reason not to use Deser.

      The second problem usually is that actually solving Serde's issues comes at a significant cost in compile time and/or runtime performance. Deser is no different. While Deser's compile times are a bit better than Serde's, the binary bloat is quite a bit worse and the runtime performance is mixed. It's roughly comparable if you look at the numbers but depending on the format structure you are losing significantly from some of the tradeoffs.

      That said, it's now in a state where it's at least in principle a drop-in replacement where the tradeoffs might work well for users.

      Deser's Design

      Deser does not try to be significantly different than Serde on the surface level. For most uses you derive Serialize and Deserialize and then start using it with your format implementing crate of choice. Most attributes are very similar, though they are taking Rust expressions instead of strings.

      use deser::{Serialize, Deserialize};
      
      #[derive(Debug, Serialize, Deserialize)]
      #[deser(rename_all = "camelCase")]
      pub struct Account {
          id: u64,
          account_holder: String,
          #[deser(default)]
          is_deactivated: bool,
      }
      
      let account: Account = deser_json::from_str(json)?;
      

      The difference in the design would become more apparent if you implement a serializer or deserializer yourself. Instead of visitors that call into each other recursively, deserializing a type creates a sink which receives events that are directly emitted by the parser, and serializing produces emitters that hand out values. Nested sinks and emitters are handed back to a driver, which keeps them on the heap. This design, which is entirely stolen from miniserde, gives some interesting consequences:

      • No stack overflows. You can arbitrarily nest structures without issues. For untrusted input you set limits with a layer, and you pick the number that you are comfortable with, which is independent of your stack space.
      • Suspendable. Because the state lives in the driver, a deserialization can be fed input as it arrives. It's also Send, so it can move between threads while you wait on IO which makes it much nicer to use with tokio. Formats like JSON, CBOR and MessagePack can be parsed as a stream if you so desire.
      • An extensible data model. The core data model is small and made of atoms, maps and sequences. For all else, there are extension values (DateTime, Uuid, etc.) that also all carry a fallback for formats that don't understand them. Unlike Serde this means it does not rely on in-band signalling of objects with magic keys to smuggle values through.
      • Lossless buffering. When a value has to be buffered (for instance because the tag of an internally tagged enum comes last), Deser records the events together with everything the format knew about them. Protocol specific extension types or error locations all survive.
      • Layers are a middleware system that sit between the format and your types and can track things like paths, enforce safety limits, rename keys or redact values without having to touch specific code paths.
      • Native flattening that doesn't buffer at all.

      On top of that are a lot of things that I just wanted to have:

      • Allow enum tags to be of any type, not just strings
      • Adapters that compose (as = Option<Vec<DisplayFromStr>>)
      • Enabling validation as an adapter
      • derive attributes that are real Rust expressions instead of strings
      • bytes as a core functionality in the data model
      • duplicate keys rejected by default and errors that point at the problem

      Here is a small configuration type that shows a few of these together:

      use deser::adapters::DisplayFromStr;
      use deser::de::Recording;
      use deser::{Deserialize, Serialize};
      use deser_encoding::Hex;
      use deser_validate::{Check, NonEmpty, Range};
      use ipnet::IpNet;
      
      #[derive(Debug, Serialize, Deserialize)]
      pub struct Config {
          // at least one 256-bit key, each written as hex
          #[deser(as = Check<NonEmpty, Vec<Hex>>)]
          secret_keys: Vec<[u8; 32]>,
          // `IpNet` knows nothing about deser, but has `FromStr` and `Display`
          #[deser(as = Option<Vec<DisplayFromStr>>)]
          allowed_networks: Option<Vec<IpNet>>,
          listeners: Vec<Listener>,
      }
      
      #[derive(Debug, Serialize, Deserialize)]
      #[deser(tag = "type", rename_all = "snake_case")]
      pub enum Listener {
          Unix { path: PathBuf },
          Tcp {
              host: IpAddr,
              #[deser(as = Check<Range<1, 65535>>)]
              port: u16,
          },
          // types this version does not know are kept and written back
          #[deser(other)]
          Other(#[deser(tag)] String, Recording),
      }
      

      Adapters are types, so Hex can go inside a Vec, and DisplayFromStr inside a Vec inside an Option. Validators are adapters too, so Check<NonEmpty, Vec<Hex>> decodes the keys and then checks that there is at least one. The catch-all variant keeps the tag and a recording of everything else in case someone wants to process it later.

      Errors are something I care a lot about, so here is what happens when a value is wrong:

      secret_keys = ["9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08"]
      allowed_networks = ["10.0.0.0/8", "fd00::/8"]
      
      [[listeners]]
      type = "unix"
      path = "/run/app.sock"
      
      [[listeners]]
      host = "127.0.0.1"
      port = 0
      type = "tcp"
      
      [[listeners]]
      type = "quic"
      host = "::1"
      alpn = ["h3"]
      
      
      
      let config: Config = deser_toml::Deserializer::from_str(input)
          .deserialize_with(|driver| driver.push_layer(PathLayer::new()))?;
      

      Note that here the tag of the internally tagged enum comes last which means that the values have to be buffered until the tag is known. In Serde this is tricky and we would lose the location if we used some tricks to add it. With Deser however, with the path layer enabled Deser you where in the structure the problem is:

      Unexpected: invalid value: must be between 1 and 65535 at line 10 column 8 (path: listeners[1].port)
      

      Deser Meta Data

      Deser really wants to be extensible, and XML is a more extreme example of the differences between Deser and Serde. Here is an Atom entry that mixes in Dublin Core for the authors:

      use chrono::{DateTime, Utc};
      use deser::Deserialize;
      use deser_value::Value;
      use deser_xml::DeserializerConfig;
      
      deser_xml::namespace!(
          atom = "http://www.w3.org/2005/Atom",
          dc = "http://purl.org/dc/elements/1.1/",
      );
      
      #[derive(Debug, Deserialize)]
      struct Entry {
          #[deser(rename = atom!("title"))]
          title: String,
          #[deser(rename = dc!("creator"))]
          creators: Vec<String>,
          #[deser(rename = atom!("updated"))]
          updated: DateTime<Utc>,
      }
      
      // entries we understand, and everything else is kept as it is
      #[derive(Debug, Deserialize)]
      #[deser(untagged)]
      enum Item {
          Entry(Entry),
          Other(Value),
      }
      
      let item: Item = DeserializerConfig::new()
          .resolve_namespaces(true)
          .from_str(r#"
              <entry xmlns="http://www.w3.org/2005/Atom"
                     xmlns:d="http://purl.org/dc/elements/1.1/">
                &lt;title&gt;Deser&lt;/title&gt;
                <d:creator>John</d:creator>
                <updated>2026-09-29T21:00:00Z</updated>
                <d:creator>Jane</d:creator>
              </entry>
          "#)?;
      

      XML uses namespaces which means that names need to be matched by their namespace, not by the prefix the document happens to use. Here the document says d: and the type says dc!. atom!("title") is just the string {http://www.w3.org/2005/Atom}title, which works because attributes are expressions. The two creators are collected into one Vec even though there is another element between them, and the text of updated goes straight into a chrono datetime. Because the enum is untagged, the entry has to be buffered before a variant is picked, and deser's buffer keeps both creators. So the result is an Entry with John and Jane.

      quick-xml, the most popular XML crate for Serde, drops the prefixes and ignores namespaces entirely, so a <x:title> from some other namespace is happily accepted as the title of the entry. The split list part though is considerably worse. A plain Entry fails with a duplicate field error for creator, unless you turn on the overlapped-lists feature (which, remember, is a global additive flag that any crate could set). That feature makes quick- xml read ahead to the end of the element and buffer everything in between, without a limit unless you set one.

      But the feature only helps when quick-xml is hooked up to the struct directly and no buffering is taking place. Wrap the struct in the untagged enum and Serde buffers the entry itself. Read from that buffer, Entry sees creator twice and fails again. The fallback is a map, which keeps only the last creator, and there is no error. With or without the feature you get this:

      Other({"creator": {"$text": "Jane"}, "title": {"$text": "Deser"}, ...})
      

      Notice how John is gone.

      Format specific extension types such as TOML datetimes are another case. TOML has them natively, Serde's data model does not, so the toml crate passes them on as a map with a magic key. In Deser a datetime is an extension value, which formats that know it keep and all others write as a string:

      let value: Value = deser_toml::from_str("released = 2026-09-29T21:00:00+02:00")?;
      
      deser_json::to_string(&value)?;
      // {"released":"2026-09-29T21:00:00+02:00"}
      deser_toml::to_string(&value)?;
      // released = 2026-09-29T21:00:00+02:00
      

      The same with serde_json::Value gives you {"released":{"$__toml_private_datetime":"2026-09-29T21:00:00+02:00"}}, and reading the value into a chrono::DateTime fails outright with invalid type: map, expected an RFC 3339 formatted date and time string.

      The Cost

      So now that you know Deser is at least in theory cool, at what cost?

      It is not free. The design relies on dynamic dispatch and on sinks and emitters that live on the heap, and that has considerable runtime overhead. In my own measurements for JSON, Deser reads somewhere between 33% faster and 60% slower than serde_json depending on the data. On average it's about 10% slower for reading. Writes are between three times as fast and 70% slower and a wash on average. For YAML and TOML it's noticeably faster than the Serde based crates, but that is more about the format implementations than the architecture.

      Compile times slightly are better, but not dramatically so. Because it doesn't monomorphize everything, release builds of derived code are about 2.3 times as fast as with Serde and that get a tiny bit better in practice for your own code as less recompilation is necessary.

      To make Deser's design work at all, it also uses unsafe internally. Most of this is to keep the chain of borrowed sinks on the heap. I feel like this is fine in the days of Miri and agents, but I know it makes some folks uneasy.

      And well, the biggest cost is that it's just not Serde.

      How Much Is There?

      Quite a lot actually which might be surprising. In addition to the core there is support for derive.

      It supports all flavorts of JSON you can think of: JSON, JSONC, JSON5 and HJSON. (Fun fact here: they are all generated out of one shared parser template) For binary handling it supports CBOR and MessagePack. Additionally it does YAML 1.1 and 1.2, TOML, XML and all three flavors of Apple's plist as well as CSV/TSV, urlencoded data and environment variables. For more crazy contraptions you can attach path info or capture location data as well as support for debug printing. You can perform validation as you parse, opt into different binary encodings in addition to base64, you can bridge to serde or capture dynamic values, transcode between formats or hook it up with tokio.

      For documentation see docs.rs/deser and the code itself is on GitHub alongside many examples.

    23. 🔗 Ampcode News Plaid Speed rss

      Amp now supports Plaid speed for modes that use GPT-6 Astra. Plaid uses OpenAI's ultrafast speed tier, which enables inference requests to run up to 6× faster, with 6× cost per token.

      Open the mode picker in the new thread dialog, choose a mode that uses GPT-6 Astra (e.g., the default High mode), and move the speed switch past Fast to Plaid.

      The mode dial with high selected and the speed switch set to Plaid, showing the red 6× cost indicator

      Plaid currently works with Amp-provided inference only, as OpenAI does not yet support ultrafast for linked ChatGPT subscriptions. When Plaid is activated, subagents and non-Plaid-enabled model inference will fall back to fast or standard speed. Like fast mode, Plaid can be toggled on or off for an existing thread.

  3. September 28, 2026
    1. 🔗 MetaBrainz Save the date: AMA with Silona Bonewald, MetaBrainz Foundation Executive Director rss

      Silona Bonewald, our new Executive Director, will be hosting an AMA (Ask Me Anything) on our forums at:

      5th October 2026
      17:00 Central European Summer Time (CEST)

      In this forum thread

      Please respect the code of conduct, as you always do!

      We are aware that this time will not suit everyone's timezone, so we will open the thread a day or so before the AMA, for early questions. We will also leave the thread open for a couple of days - but don't wait too long!

      You can find more information about Silona on our Executive Director announcement post.

    2. 🔗 Quarkslab's blog From AI Agents to RCE - Building a Vulnerability Research Workflow rss

      Introduction

      A security researcher spends most of the day reading code. Thousands of lines, sometimes, just to find the few that matter. Most of that reading is not where the meaningful work happens. The real work is the moment of suspicion : this length field is validated here but not there, this state can be entered twice, this loop trusts a value it should not trust. That moment takes seconds. Reaching it can take days.

      An agent can work across a codebase in minutes, connecting functions scattered across files and following relationships that would take a researcher much longer to piece together by hand. What is less clear is what comes out of that volume: real vulnerabilities, or a pile of false positives to sort through.

      We wanted to turn this raw power into a method. We built an agent harness for vulnerability research: the surrounding system that manages the tools, context, structure, and rules AI agents need to work on a task. In our case, that means a queryable graph over the target, a deterministic analysis core, a set of lenses that read its output, and a staged pipeline the agents work through. The harness proposes attack vectors, explores the code paths involved, and writes proofs of concept, with a researcher checking every step that matters.

      The goal was not a push-button vulnerability finder, but a methodology that can be run, inspected, and repeated : the agents handle the groundwork, correlate what they find, and bring up leads, while the researcher validates them, directs the investigation, and decides what deserves more time. What we were really after was a researcher who spends the day suspecting rather than scrolling.

      This article follows the workflow from the first pass over a codebase to a working exploit, showing the output of each stage along the way. We tested it on FreeRDP10, an open-source RDP client, a large target, where it uncovered two vulnerabilities that can be chained into remote code execution on the client.

      The problem, and what we did about it

      Late in December 2025, one person targeted nine Mexican government agencies1. Not a crew, not a state program. Over the following seven weeks they typed 1,088 instructions , which produced 5,317 commands executed across 34 separate sessions. Claude Code handled around 75% of the live exploitation across 305 internal servers , while a second pipeline built on GPT-4.1 read the takings and tasked the next move. Close to 200 million records left the building: tax filings, the civil registry, patient data, vehicle records, the electoral roll. The forensics come from the attacker's own recovered servers.

      The useful number here is the ratio: roughly five commands were executed for every instruction typed. A similar pattern appeared three months earlier in a China-linked campaign2, where the model reportedly handled 80 to 90 percent of the tactical work against roughly thirty organisations.

      For defenders, the important change is the amount of work one operator can sustain. The Mexico operation was run by one person , while automation covered tasks that would previously have required several skill sets. Defenders often assume that an attacker's skills set a ceiling on what they can do; agentic tooling weakens that assumption.

      The campaign did not rely on novel exploitation techniques. Its significance was the scale at which one operator could carry out the work. That raises the defensive question behind this project: if automation can reduce the cost of exploiting known weaknesses at scale, can it also reduce the cost of finding and understanding new ones? We built the workflow to test that question, and pointed it at FreeRDP.

      That same ability to produce work at scale creates a different problem for defenders: more output does not necessarily mean more useful findings. curl illustrates this well34. Its confirmed-vulnerability rate fell from above 15 percent of submissions to below 5 percent, with about one report in five during 2025 amounting to machine-generated noise, each still costing hours of a volunteer team's time. The programme eventually closed. The reports were costly because many were plausible: they cited real functions and code paths and described attack scenarios that required investigation before they could be dismissed.

      The limiting resource is therefore researcher attention , not the number of findings a system can produce. We designed the workflow to search broadly, then spend machine time filtering and validating candidates before asking a researcher to inspect them.

      What already exists, and why we built our own anyway

      Several existing systems address adjacent parts of this problem, using different approaches: Google's Big Sleep for AI-assisted vulnerability research, OpenAI's Aardvark for continuous security analysis of code repositories, and the autonomous systems developed for DARPA's AI Cyber Challenge (AIxCC).

      Their results show that agent-assisted vulnerability research can work. Big Sleep56 turned up a memory-corruption bug in SQLite that threat actors already knew about, and reported twenty more in projects such as FFmpeg and ImageMagick. Aardvark7, now Codex Security8, builds a threat model from a repository, tests each commit against it, and validates what it finds in a sandbox. DARPA's AI Cyber Challenge9 took a different approach, with systems designed to find and patch vulnerabilities autonomously.

      Each was built for a job adjacent to ours, not for ours. Aardvark7 watches a repository as you commit to it, which is the right instinct when the code is yours and no help at all when you are auditing something you have never touched. Big Sleep5 is an internal Google tool. The AIxCC9 systems presume a fuzzing harness already exists and are scored on running unattended; the FreeRDP surface we cared about had no harness, and we had no intention of running unattended.

      The deeper reason is simpler. A tool that decides everything by itself hands you a conclusion and nothing else. You then have to redo the work to find out whether it is true, which costs roughly what finding it would have cost. curl's maintainers spent 2025 doing that with reports from strangers. We had no intention of doing it with reports from our own tooling.

      The workflow, in one picture

      The architecture separates two properties we wanted to use differently: language models are useful for generating and exploring hypotheses, but their outputs are not reproducible (due to their non-deterministic nature).

      The system therefore has two layers. A deterministic core indexes the code into a queryable graph, runs static taint from attacker-controlled sources to dangerous sinks, and passes the result through a set of narrower lenses: known CVE and CWE patterns, and a handful of checks specific to the target. An agent layer above it forms hypotheses, reads what the core has surfaced, judges what is real, and drives it towards a PoC.

      The core is tuned for recall rather than precision. It over-produces on purpose, because a lead we never generate is one we can never recover, and the agent layer exists to make that over-production affordable. It is not there to find bugs. It is there to absorb the noise a broad search inevitably produces , so that a researcher sees only what survived.

      The core is reproducible. The agent layer is not, and does not need to be. It needs to be auditable.

      Concretely, the pipeline runs in eight stages:

      graph -> recon -> slicing -> analysis -> triage -> poc -> chain -> exploit
      

      Stage | What it does
      ---|---
      graph | Indexes the codebase.
      recon | Turns structure into attack surface.
      slicing | Narrows things to a single hypothesis - no agent reasons well across an entire repository.
      analysis | Runs taint, pattern matching and the other lenses over that slice.
      triage | Sorts the output into real , unclear and noise.
      poc | Attempts to make a survivor fire.
      chain | Asks whether two weak primitives amount to one strong one.
      exploit | Turns a validated primitive or chain into a working exploit.

      The eight-stage
workflow

      The researcher works between the stages, never inside them. Which also means they can enter at any of them: run the whole thing end to end, stop after recon and take the hypotheses elsewhere, or hand the pipeline a finding of their own and let it do the triage and the poc. No stage assumes the one before it was run by us rather than by a person.

      The rest of the article applies these stages to FreeRDP and shows the artefacts produced at each step.

      Dive into the workflow through the FreeRDP example

      We pointed the workflow at the FreeRDP10 tree at commit 993499447 and gave it nothing else to go on. No list of past FreeRDP CVEs, no fuzzing corpus, no hint about which subsystems have historically been soft. git clone, and go.

      FreeRDP is the library most Linux RDP clients are built on: Remmina, GNOME Connections, KRDC, and the RDP support in Apache Guacamole's guacd. The same tree also implements the server side, which is what GNOME Remote Desktop and KRDP use, so parsing code shared between the two roles is exposed from both directions. We audited the client, where every byte parsed from the network ultimately comes from a server the user has chosen to trust.

      Step 1 - graph, recon and slicing

      graph (*) -> recon (*) -> slicing (*) -> analysis -> triage -> poc -> chain -> exploit
      (*) Steps explained below
      

      Graph indexes the target into a representation the agents can query: 13,506 functions, 3,105 types and 31,705 call edges in this run, including synthetic edges for resolved function pointers.

      The graph is not meant to be explored as a whole. Its value is in narrowing the view. libfreerdp/core/nego.c, for example, handles the negotiation exchange at the beginning of an RDP connection. Resolving that file gives the agent its definitions, parameters and relationships to the rest of the codebase.

      Figure - nego.c resolved in the
graph.

      From there, reachability queries turn an exposed entry point into a working set:

      $ sift graph reachable nego_recv --files
      libfreerdp/core/nego.c
      libfreerdp/core/tpdu.c
      libfreerdp/core/tpkt.c
      … 6 more
      

      Starting from nego_recv, the working set comes to nine files. Other entry points are broader: rdpgfx_recv_pdu reaches 67 files and rdp_recv_pdu reaches 65. Slicing uses those relationships to keep each investigation small enough for an agent to reason about without carrying the whole repository in context.

      Recon adds exposure information. It groups the code into components, estimates how directly attacker-controlled data reaches each one and produces a threat model, a barrier map and an ordered slice list.

      Figure - Components by
exposure.

      The run produced 304 slices. The two findings discussed below came from slice 166, libfreerdp/core/nego, and slice 109, channels/urbdrc/client/data_transfer. A purely mechanical ranking would not have selected either one early. The researcher used the ranking as a map, not as a verdict.

      Step 2 - analysis

      graph -> recon -> slicing -> analysis (*) -> triage -> poc -> chain -> exploit
      

      Each slice is analysed independently. The deterministic core builds a deeper graph, runs static taint and applies the configured lenses. Taint runs twice: first with sources and sinks already known to the engine, then again after an agent has identified target-specific vocabulary such as FreeRDP stream operations, read macros and allocation wrappers.

      In slice 166 the engine produced 14 findings. Two became interesting together:

      f-007  attacker-controlled length stored in nego->RoutingTokenLength
                               &varr ?
      f-005  same field later used as a Stream_Write size
      

      The engine could establish both endpoints but not the relationship across the lifetime of the nego object. It reported them separately instead of inventing a data-flow edge. This is the kind of ambiguity the next stage is meant to resolve.

      Figure - f-005 and f-007
combined.

      Slice 109 produced 116 findings. Two pointed to a USB-redirection completion path where an error leaves an attacker-supplied OutputBufferSize in place even though no data has been written. The client later uses that size when emitting the response, exposing bytes left in the heap allocation.

      At this point, the two slices have produced 130 findings, but these are still candidates, not confirmed vulnerabilities. False positives are expected at this stage, and each finding still needs to be investigated and validated.

      Step 3 - triage and PoC

      graph -> recon -> slicing -> analysis -> triage (*) -> poc (*) -> chain -> exploit
      

      Triage groups related findings before classification. In this run, 130 findings became 36 clusters: 20 real and reachable, 3 latent, 3 mixed and 10 false positives.

      The classifications are then checked dynamically. For each candidate, the pipeline builds a small harness against FreeRDP and runs it under AddressSanitizer. This is where a plausible static finding either becomes measurable or gets discarded.

      For slice 166, triage brings f-007 and f-005 back together. f-007 ends at nego_set_routing_token, where the length derived from network data is stored in nego->RoutingTokenLength. f-005 picks up the same field later, when it is passed as the length argument to Stream_Write in nego_send_negotiation_request.

      Connecting the two findings establishes a candidate path, but not yet a vulnerability. There are two ways for nego->RoutingTokenLength to be populated. The first turned out to be safe: a protocol length field limits the routing token to a size that cannot overflow the destination buffer. The PoC confirmed the bound, so that path was discarded.

      The second path comes from a Server Redirection PDU. A malicious server can provide LoadBalanceInfo, which is later reused as the routing token when FreeRDP establishes the redirected connection. This path does not have the same length restriction.

      We reproduced it using a malicious server and an ASan-instrumented FreeRDP client. The server sent a 600-byte LoadBalanceInfo payload, which became a 615-byte routing token once wrapped with the cookie header. At Stream_Write, FreeRDP attempted to copy those 615 bytes into the 512-byte allocation created by nego_send_negotiation_request:

      ==4514==ERROR: AddressSanitizer: heap-buffer-overflow
      WRITE of size 615
      
      #2  Stream_Write
      #4  nego_send_negotiation_request  libfreerdp/core/nego.c:1098
      #9  rdp_client_redirect            libfreerdp/core/connection.c:715
      
      0x75e976227180 is located 0 bytes after 512-byte region
      [0x75e976226f80,0x75e976227180) allocated by thread T2 here:
      
      #1  Stream_New
      #2  nego_send_negotiation_request  libfreerdp/core/nego.c:1083
      

      The rdp_client_redirect frame confirms that the overflow is reached through the server-driven redirection path rather than by calling the vulnerable function directly from the harness.

      The test used WITH_VERBOSE_WINPR_ASSERT=OFF, the configuration used by Debian and Ubuntu builds. With verbose assertions enabled, the bounds check aborts the client before the out-of-bounds write reaches memcpy.

      The USB-redirection finding went through the same validation process. In that case, the candidate primitive was an uninitialised heap disclosure. We configured ASan to fill new heap allocations with 0xcd; the same pattern appeared unchanged in the PDU emitted by FreeRDP:

      [poc] heap pattern bytes (0xcd) in region [36..100): 64 / 64
      

      Only findings that survived this dynamic validation were passed to the chaining stage.

      Step 4 - chaining

      graph -> recon -> slicing -> analysis -> triage -> poc -> chain (*) -> exploit
      

      The chaining stage works only on validated clusters. Here, the negotiation bug provides an attacker-controlled but blind heap write, while the USB- redirection bug provides visibility into heap contents. The leak operates after the session is established, and a server redirection can then force the client to reconnect and trigger the write. Together, these primitives made remote code execution plausible, but not yet demonstrated.

      Step 5 - exploitation

      graph -> recon -> slicing -> analysis -> triage -> poc -> chain -> exploit (*)
      

      Exploitation changed the balance between the agent and the researcher. Earlier stages produce artefacts that are cheap to check: a path, a source location, an ASan trace or a measurement. Exploitation is less forgiving. A failed heap- layout hypothesis, for example, often says little about why it failed.

      So we gave the exploitation stage its own harness rather than letting a general-purpose agent iterate freely. It provides exploitation-specific knowledge and breaks the process into a set of gates that must be validated before the agent can move on.

      This kept the agent useful on bounded tasks: interpreting leaked bytes, enumerating objects reachable from the write, testing heap-layout assumptions and generating harnesses for individual experiments. It also made its progress inspectable. A successful step meant that an expected property had been measured, not simply that the model considered the approach plausible.

      There were still limits. During this run, the agent did not converge on the complete exploit by itself. We intervened twice when the models we used stopped making useful progress: first while deriving the libc base from the leak, and later while composing the primitives into a working payload.

      The exploitation proceeded through the following stages:

      ✔  Lab set up: client VM, hostile server
      ✔  Cluster 2 write primitive measured
      ✔  Control-flow hijack demonstrated
      ✘  libc base from the cluster 28 leak
      &boxh&boxh researcher intervention 1 &boxh&boxh
      ✔  libc base recovered
      ✔  ASLR defeated
      ✘  Working payload
      &boxh&boxh researcher intervention 2 &boxh&boxh
      ✔  Leak and write composed into one session
      ->  /tmp/pwned
      

      RCE attempts on the client
side

      Conclusion

      On FreeRDP, the workflow produced 130 findings across the two slices discussed in this article. Triage reduced them to 36 clusters, 20 of which were reproduced as real and reachable. Two findings provided the primitives used for client-side remote code execution.

      The workflow did not eliminate noise, nor was it designed to. The deterministic core could search broadly because the researcher was not expected to inspect every finding it produced. Slicing kept individual investigations bounded, clustering brought related evidence together, and PoCs provided a way to reject or confirm the resulting hypotheses before more time was spent on them.

      The slice ranking also showed why we kept a researcher in the loop. The two relevant slices were ranked 166th and 109th out of 304. The ranking could account for properties such as parsing volume and exposure, but not every reason a component might be interesting to a security researcher. It provided an ordered search space rather than deciding what should be investigated.

      From git clone to the report sent to the maintainers took four days on a roughly half-million-line C codebase that has been audited and fuzzed for years. The model was only one part of that process. Around it were the pieces that made its output usable for vulnerability research: a queryable representation of the code, deterministic analysis, bounded slices, evidence attached to findings, dynamic validation, an exploitation harness, and researcher checkpoints.

      Both vulnerabilities were reported to the FreeRDP maintainers. FreeRDP was our test case; the objective was the workflow itself, and whether it could make a large codebase practical to investigate without turning the researcher's time into the bottleneck.

      Advisories :


      References


      1. The AI-Assisted Breach of Mexico's Government Infrastructure (Eyal Sela, Gambit Security, technical report, April 2026) &larrhk

      2. Disrupting the first reported AI-orchestrated cyber espionage campaign (Anthropic, November 2025) &larrhk

      3. The end of the curl bug-bounty (Daniel Stenberg, January 2026) &larrhk

      4. Death by a thousand slops (Daniel Stenberg, July 2025) &larrhk

      5. From Naptime to Big Sleep (Google Project Zero, November 2024) &larrhk&larrhk

      6. Google's AI security announcements (CVE-2025-6965, the SQLite flaw known to threat actors) &larrhk

      7. Introducing Aardvark: OpenAI's agentic security researcher (OpenAI, October 2025) &larrhk&larrhk

      8. Codex Security: now in research preview (OpenAI, March 2026) &larrhk

      9. AI Cyber Challenge marks pivotal inflection point for cyber defense (DARPA, August 2025) &larrhk&larrhk

      10. FreeRDP sources (GitHub) &larrhk&larrhk

    3. 🔗 HexRaysSA/plugin-repository commits update danielplohmann/mcrit-plugin to familiary/mcrit-plugin (reposit… rss
      update danielplohmann/mcrit-plugin to familiary/mcrit-plugin (repository transferred)
      
    4. 🔗 exe.dev Pools: Our Newest Primitive rss

      At exe, one of the things we pride ourselves on is developing strong primitives instead of throwing features at the wall. We still believe that the VM is a key primitive in our system. Everything is just a computer and you get to choose how you use it.

      We’ve encouraged you all to build VMs, scale them, and run your production stacks on our platform. But we noticed we were missing a primitive for logically grouping VMs. Today, we’re launching pools.

      A pool is a collection of VMs that share a set of resources. Before pools, a VM’s resources came directly from your plan. You could scale VMs independently, but you couldn’t define a production or staging stack with its own resources in its own region.

      Pools also give you predictable pricing. We don’t want to send you a surprise invoice because one VM had a burst of traffic.

      Here’s how pools work:

      1. Create a pool in a region, with a specific CPU and memory size.
      2. Move existing VMs into the pool, or create new VMs in it.
      3. Those VMs now share the pool’s resources.
      4. Set up more pools as you need them.

      Pricing

      Pools come with some plan changes. New customers can start using pools today, and existing customers can migrate to the new plans to get access. There are two new plans to choose from: Personal and Work.

      Personal

      The Personal plan starts at $15 a month for a 2 vCPU / 4 GB pool that holds up to 50 VMs. You can scale your pool up to 16 vCPUs / 32 GB of memory for $155 a month. Additional disk and bandwidth charges still apply.

      We’re also introducing standalone VMs, billed at an hourly rate. Other platforms call these sandbox VMs. Customers have been asking for this for a while, and we wanted to make sure we built the right thing. You manage a standalone VM like any other VM, but you own its lifecycle and decide how to use it. You can think of it as a VM in a pool of 1.

      Work

      Our previous Teams plan charged per seat. This made sense at first, but customers told us they’d rather pay for compute. So the new Work plan no longer charges for seats.

      The Work plan is built around pools. It has a minimum spend of $150 a month, and the first 4 vCPUs / 8 GB of memory are included. On top of that, you can create as many pools as you want, of any size, billed at a predictable rate starting at $0.105/hr per 2 vCPUs. Like the Personal plan, it also includes standalone VMs.

      Companies come in all shapes and sizes. If the Work plan doesn’t fit your team, reach out and we’ll work with you to find a better solution.

      Pricing changes can be hard, but we believe pools are a strong foundation moving forward. Existing customers can take their time moving over (a year, to be exact). We’re here to help, and we’re excited to see what you build.

    5. 🔗 hacker news ida pro references New comment by jchw in "Sonnet 5.5" rss

      I have been trying to convince the safe guards that analyzing a C++ compiler from 2003 isn't particularly relevant to modern cybersecurity. It seems Anthropic disagrees.

      IDA Pro and Ghidra, thankfully, still lack such safeguards...

      (No other model I've tried has refused either FWIW.)

    6. 🔗 anthropics/claude-code v2.1.284 release

      What's changed

      • Added Claude Sonnet 5.5 (claude-sonnet-5-5), now the default Sonnet model on the Anthropic API — 1M context, $2/$10 per Mtok with $0.20/Mtok cache reads
      • Added a "Yes, but ask again next time" answer to auto mode's prompt before a read outside the working directories, so you can allow that one read and still be asked about later ones
      • Added dollar amounts to the Claude apps gateway spend limit in /usage and the status line (for example "$271.40 / $500.00 spent this month") when the gateway runs this version or later; the status line's rate_limits.spend_limit also gains used_usd, limit_usd and period
      • Added effortSlider:decreaseEffort, increaseEffort and toggleUltracode keybinding actions, so the /effort slider's arrow and Tab keys can be rebound in keybindings.json
      • Added /rate-limit-options to /help and the command menu for claude.ai subscribers, so the usage-limit notices that mention it point to a command you can find
      • Added /mcp reconnect all in the interactive terminal to retry every MCP server that failed to connect or needs authentication at once
      • Added Claude apps gateway startup warnings when a managed policy's availableModels is empty, or leaves out the model Claude Code starts on without setting model or enforceAvailableModels
      • Added auth: { google: {} } for Claude apps gateway telemetry.forward_to destinations, so telemetry can be exported straight to Google Cloud's OTLP endpoint using the gateway's Google Cloud credentials
      • Added certificate client authentication (private_key_jwt) between the Claude apps gateway and its identity provider, for identity providers that issue certificate credentials instead of client secrets
      • Fixed a damaged response stream showing raw errors such as "JSON Parse error" or "undefined is not an object", or writing the word "undefined" into an answer, instead of being retried or reported as an interrupted response
      • Fixed an overloaded or server error arriving right after a thinking block ending the turn with an error instead of being retried
      • Fixed "Prompt is too long" errors that persisted after compacting: when the compacted request is still too long, Claude Code now compacts once more, keeping less of the recent conversation
      • Fixed a session whose model is unavailable, with no fallback model left, showing a bare "is currently unavailable" message (or "Something went wrong" in cloud sessions) instead of the model-unavailable notice and its Learn more link
      • Fixed Agent SDK sessions crashing when a user message contains an image with a malformed source, and failing on every later turn after a malformed document block; a malformed image is now replaced with an explanatory note
      • Fixed MCP tool calls in a resumed session failing with "No such tool available" while their server was still connecting; the call now waits up to 10 seconds for the server
      • Fixed repeated calls to the plan-usage endpoint after it rate-limits or rejects your login: /usage, /extra-usage and IDE usage views now back off instead of re-asking
      • Fixed claude mcp add reporting success when managed settings restrict MCP servers to plugins; it now refuses and says what to do, instead of saving a server that never loads
      • Fixed the /plugin configure screen: boolean options are now a true/false choice instead of free text, number options refuse invalid input, and ←/→ change an options field instead of switching tabs
      • Fixed ANTHROPIC_FOUNDRY_RESOURCE being interpolated into the Foundry endpoint host unvalidated; a value that is not a plain resource name is now refused
      • Fixed Claude Desktop behind a Claude apps gateway offering no 1M context option: the gateway now marks each 1M-capable model for Desktop automatically
      • Fixed ↓ in shell mode selecting a hidden background-tasks pill, which stopped Backspace and Ctrl+U from editing the prompt
      • Fixed Bash tool failing on Windows with many plugins enabled: plugin bin/ directories that don't exist are no longer added to PATH, and inherited entries aren't added twice
      • Fixed sparsePaths plugin marketplaces cloning empty and replacing a working local copy on older git (before 2.39), which failed every refresh with "marketplace.json file is no longer present"
      • Fixed fullscreen rendering erasing the terminal output above the session when [ in transcript mode writes the conversation to scrollback (macOS and Linux)
      • Fixed fullscreen scroll position jumping to the previous message or to the bottom when a reply finished streaming while scrolled up
      • Fixed tab bars in dialogs such as /config and /plugin breaking the title and tab labels mid-word in a narrow terminal; a tab that doesn't fit now moves to the next line whole
      • Fixed the /model picker showing "+1 model" below the list after scrolling to the last model; the count now covers only the models below the visible rows
      • Fixed /keybindings writing Backspace and Delete bindings for a footer action that does nothing into the generated keybindings.json
      • Fixed a rebound agent panel close key (footer:close) typing "x" instead of itself on the row of the agent you're viewing
      • Fixed vim mode . not repeating text typed very fast (for example over ssh or in tmux) or pasted without bracketed paste, and leaving the prompt in INSERT mode after repeating a change with nothing typed (such as cw then Esc)
      • Fixed vim mode leaving the cursor on an image placeholder's opening bracket after dd on the last line or yy at the end of the prompt, where r or x would break or delete the image
      • Fixed a key pressed the instant the terminal regained focus answering the Remote Control enable prompt before its short safety delay restarted
      • Fixed the workspace trust dialog appearing a second time after switching renderers or updating when Claude Code was started in the home directory
      • Fixed rules symlinked into .claude/rules from outside the project being skipped without ever showing the external-imports approval prompt; a .claude directory symlinked from outside the project now asks for the same approval
      • Fixed plugins from marketplaces, claude.ai and npm pre-approving their own tools via allowed-tools under managed allowManagedPermissionRulesOnly; only plugins from an official Anthropic source or a source that managed settings vouch for keep that pre-approval
      • Fixed a failed first claude plugin install leaving the plugin enabled and recorded when a dependency's version range could not be met
      • Fixed the debug log dropping a failed hook's stderr when the hook also wrote to stdout, and logging nothing for a failed hook with no output; failed hooks now also log their status code
      • Fixed {"decision":"block"} returned by Elicitation and ElicitationResult hooks being ignored; it now declines the MCP elicitation, as exit code 2 does
      • Fixed sessions launched without the SendMessage tool (such as by Claude Desktop) still being told to message other sessions with it
      • Fixed a photo sent from the Claude app over Remote Control being lost when its queued message was pulled back into the terminal prompt to edit, and the cursor moving one character for a photo with no caption
      • Fixed typing a message during an automatic usage-limit wait taking the wait out of the "Continue automatically at usage limit" setting's control when that turn hit the limit again
      • Fixed usage-limit warnings suggesting /upgrade to users already on the highest Max plan; the warnings and /upgrade itself now point at /usage-credits when it is available
      • Fixed the Explore subagent switching to Opus on the Claude API when the session runs a model ID Claude Code doesn't recognize, such as a custom model behind a proxy; Explore now inherits that model
      • Fixed /loop status updates in self-paced mode often not being shown because Claude wrote them only in its reasoning; Claude now writes each update, and the outcome when the loop stops, as visible text
      • Fixed /ultrareview failing to upload the working tree when started from a git worktree that the Claude desktop app created on macOS or Linux
      • Fixed sandboxed Bash commands failing to start on Linux when the working directory is write-denied and contains a read-denied directory
      • Fixed artifact database write results telling Claude that every viewer sees a write to a viewer's private data/users/ subtree, and added a "view" level to as_level
      • Fixed the Claude apps gateway answering 431 Request Header Fields Too Large to every request from a sign-in whose identity provider lists many groups; it now accepts request headers up to 256 KiB
      • Improved the usage-limit wait: the limit's state and the countdown with the usage-credits option now show as one block under the prompt, and limit messages no longer repeat the countdown
      • Improved the "No such tool available" error for Claude in Chrome tools called without their prefix: it now names the tool to call
      • Improved Monitor event rows to show what each event printed instead of repeating the description, and stopped repeating an unchanged "Waiting for N … to finish" line after every event
      • Improved Workflow tool sandbox hardening for errors thrown by async script hooks
      • Improved startup time and memory use by building only the parts of the settings schema that your settings files actually use
      • Improved /claude-api: hillclimb no longer spends rounds on prompt rewordings too small for the eval to measure, and an extra page you ask for beside report.html is built as one local file that loads nothing from the network
      • Improved lists such as /tasks, /copy and /hooks: the details after each name now line up in one column when they fit, and otherwise sit at the right edge
      • Improved claude plugin marketplace add to say when it replaces a marketplace already added under the same name from a different source, and how to undo it
      • Improved the startup refusal when managed settings require a sign-in (forceLoginMethod or forceLoginOrgUUID) and an API key, token or apiKeyHelper is configured: it now names the credential in use, where it is set, and how to remove it
      • Improved auto-memory loading: invisible characters and tags that imitate Claude Code's own markup are neutralized in MEMORY.md and recalled memory notes before they reach Claude
      • Improved claude remote-control: in a folder you haven't trusted yet, it now asks for workspace trust on the terminal instead of exiting
      • Improved artifact pages: Claude writes its design plan into the page instead of the reply, and uses the name you already gave something as the page title
      • Improved the Artifact tool so that when Claude is given a claude.ai chat or project link, an artifact from a chat, or an artifact id on its own, it asks for the right link or the content instead of stopping
      • Changed interactive terminal and VS Code sessions to start in auto mode when no permission mode is configured, on every plan and provider; permissions.defaultMode still overrides it
      • Changed Ultracode into its own toggle in /effort (Tab, or /effort ultracode [on|off]): it no longer forces xhigh effort and stays on at any effort level
      • Changed retries after a dropped connection mid-response to share one budget with the rest of the request's retries, so a failing request gives up sooner
      • Changed the notice shown when a Sonnet model's safeguards flag a message to explain why it happened and to offer editing and retrying
      • Changed safety-related model switches in sessions that pin an Opus model with ANTHROPIC_DEFAULT_OPUS_MODEL or modelOverrides: on the Anthropic API, the API now picks the model to switch to for each kind of flag, not the pinned model
      • Changed the non-interactive first turn to still wait up to 2s for connecting MCP servers named by --allowedTools or an mcp_tool hook, even when CLAUDE_CODE_MCP_STARTUP_WAIT_MS is 0
      • Changed /recap to decline with a short notice when it arrives relayed from a chat thread (your own included) or from a routine or webhook; typed in the terminal, the Claude apps, Remote Control, -p or an SDK host, it runs as before
      • Changed /artifacts to show its filter tabs beside the title with one-word labels (All, Mine, Shared), using the same tab bar as /config and /plugin
      • Changed artifact publishing to refuse a file on a network share (a \\host\share path or a /net automount) unless it is on a mapped network drive added with --add-dir
      • [VSCode] Added an optional time above each prompt and response, with a date line where the day changes (Claude Code: Show Message Timestamps setting, off by default)
      • [VSCode] Added plugin load errors and notes to the Manage plugins rows, with a popup to disable, uninstall or copy the error
      • [VSCode] Added an Ultracode on/off switch under the Effort slider, replacing the slider's Ultracode stop; the model pill shows "· Ultracode" at any effort level
      • [VSCode] Fixed Reload Claude from the Memory dialog restarting before an edited file was saved
      • [VSCode] Fixed a restored tab opening a conversation another Claude process still has open; it now asks first
      • [VSCode] Fixed Focus view sections you expanded closing on their own while a sub-agent is working or when the section's first step is trimmed from view
      • [VSCode] Fixed typing /model and Enter printing usage text into the chat instead of opening the model selector
      • [VSCode] Fixed /feedback on Vertex, Bedrock and Foundry being refused after you pressed Send; the report is now saved on this computer, as the terminal does
      • [VSCode] Fixed sign-in waiting up to a minute for the Python extension after a window reload
      • [VSCode] Fixed Claude Code tabs that stopped responding after Restart Extensions: they now reopen on their conversation
      • [VSCode] Fixed a message from another agent with no recorded sender showing as raw XML in the chat
      • [VSCode] Fixed messages from other agents, sessions or channels disappearing after a reload
      • [VSCode] Fixed a user's own /mcp, /config or /settings command being shadowed by the extension's dialog
      • [VSCode] Fixed Escape stopping every background agent when no turn was running
      • [VSCode] Fixed plugin install links replacing a marketplace you already have that uses the same name
      • [VSCode] Fixed "Prompt is too long" errors after compaction when a large text file is attached to a message
      • [VSCode] Fixed chat links to files with non-ASCII characters, spaces or brackets in their path not opening
      • [VSCode] Changed CLAUDE_CONFIG_DIR in the claudeCode.environmentVariables setting to apply only when it is an absolute path, and passed it to terminals that continue the chat
      • [Cloud sessions] Fixed a routine's Edit and Duplicate controls saying the routine was still loading while you were offline; they now tell you you're offline
      • [Claude Tag] Added model family choices such as "Opus (latest)" for a thread, a channel default or your DM, so the choice follows the newest model in that family
      • [Claude Tag] Added the spend that counts toward your organization-wide limit to the analytics spend projection chart, with how much of the limit is used
      • [Claude Tag] Fixed the earlier Claude in Slack app's progress card and link previews omitting the repository and Create PR button when a GitHub Enterprise host name contains an underscore
      • [Claude Tag] Fixed Claude staying silent in a channel whose environment declines to start it; it now posts one notice asking you to contact an admin, and retries when @-mentioned
      • [Claude Tag] Changed Claude to post its private sign-in notice at every @mention from someone who hasn't connected their Claude account, instead of going quiet after the first
      • [Claude Tag] Improved "Notify members now" in admin settings: one press reaches every workspace your organization claimed in an Enterprise Grid, and more members in large workspaces
      • [Claude Tag] Improved Claude's wait notice on self-hosted environments with on-demand runners: it now says whether a runner is starting, a start will be retried, or no runner will start
      • [Claude Tag] Improved the error shown when adding a channel manager fails because the channel's Slack workspace can't be confirmed as connected to your organization
      • [Claude Tag] Improved a channel's access lists in admin settings to show the connectors, repositories and plugins an auto-join pattern attaches, and where each comes from
      • [Claude Tag] Improved adding repositories as a channel manager: when your GitHub sign-in can't confirm you're a repository admin, the page asks you to sign in with GitHub
      • [Code Review] Fixed Code Review giving up without posting a finished review when an unsubmitted review under its GitHub App was open on the pull request; it now retries the post first
    7. 🔗 MetaBrainz GSoC 2026: GraphQL Server For Musicbrainz rss

      Hi everyone, I'm Sreehari, also known online as owlpharoah (op3kay on Matrix). I'm a second year student at IIIT Jabalpur. This summer I worked on the foundations of a GraphQL server in Rust that sits over the MusicBrainz PostgreSQL database, under the guidance of @bitmap and @jadedblueeyes.

      The Setting

      MusicBrainz already has an XML/JSON API, but getting related data out of it means chaining together inc parameters, and browsing support differs from one entity type to the next. Looking up five artists at once isn't really possible either.

      GraphQL fixes most of this by letting a client ask for exactly the fields and relationships it wants in one query. It also lets the server check how expensive a query is before running it, instead of finding out after the database has already taken the hit.

      Before writing the proposal, I built a rough prototype covering Artist, Release Group, Release, and Recording, mainly to see how things would click together. Two problems showed up: N+1 queries on relationship fields, and the fact that depth limiting alone doesn't catch a shallow query that's still expensive. Both became core parts of the proposal.

      The Plan

      The proposal scoped the project to six entity types: Artist, Release Group, Release, Recording, Label, and Area. Each would be queryable by MBID, with the usual relationships between them, plus aliases, tags, genres, and ratings across the board.

      A few decisions were made initially:

      • DataLoaders would follow a two tier split. One loader maps an entity's internal id to the hydrated entity itself and gets reused everywhere that entity shows up. A separate, thinner loader maps a parent id to a list of child ids.
      • Fields that need a loader call would live behind ComplexObject, so they only run when a client actually asks for them.
      • Pagination would use keyset pagination instead of offset pagination, since offset pagination gets slow and inconsistent on large tables that change often.
      • Query safety would come from depth limiting plus a complexity weight on every resolver.

      key goals

      • A working GraphQL server covering the six entity types
      • Schema level depth limiting and query cost analysis
      • A performance baseline from load testing.

      The Result

      DataLoader infrastructure is in place across all six entities, following the two tier split. Loaders exist for tags, ratings, artist credit, genres, annotations, aliases, ISNI and IPI identifiers, and MBID to internal id resolution. Hydration loaders are shared across every relationship that points at a given entity instead of duplicated per relationship.

      MBID redirect handling lives inside each loader's load function. When a primary table lookup misses, the unresolved MBIDs get batch queried against the matching *_gid_redirect table, and any hits get merged into the result map. A resolver further up never has to know a redirect happened.

      Keyset pagination runs across the paginated fields using ROW_NUMBER() OVER (PARTITION BY parent_id ORDER BY child_id) in a single batched query, so one query can apply a per parent limit across a whole batch of parents. The cursor ended up as a plain integer rather than the opaque string from the proposal.

      Query complexity weights reflect actual database work: a scalar field costs its default, a single hop DataLoader field costs a flat amount, a paginated one to many field scales with the requested page size, and multi hop fields carry a multiplier on top.

      Integration tests cover all six entities.

      Week 7's load testing with k6 turned up two findings worth fixing. There was an N+1 on the isrc field, fixed with a dedicated RecordingIsrcLoader, and a gap in the complexity limiter where a pathological query executed instead of getting rejected outright.

      On the infrastructure side, CI runs pre commit hooks and no longer has dead code warnings, and documentation is wired up through Magidoc in Docker compose.

      What I Learned

      The two tier loader split sounded simple on paper, but it's saved me a lot of time in practice. Adding a new relationship is now mostly copying hydration logic that already works, instead of writing it fresh.

      Fixture data needs to be checked against the database it's running against, not assumed to be stable. musicbrainz-docker's sample dumps import non- deterministic subsets of the data, so an MBID that resolves cleanly on my machine can point at something else, or nothing, on someone else's. That cost me a few confused debugging sessions before I figured out what was going on.

      What's Next

      • Criterion benchmarking, to compare the current per row queries on ComplexObject fields like Release.date against a batched loader variant, and to compare the two hop ArtistCredit and Tags loader patterns against a single joined loader.
      • Moka caching is still an open evaluation.
      • Extended entity coverage beyond the original six.

      Conclusion

      This was an awesome summer and i enjoyed thinking about and wiring up the API schema and the Postgresql database, fixing bugs, and everything in between. It was really satisfying watching the two tier loader pattern click into place once and then just work for every relationship added after it.

      Thanks to my mentors for the guidance, especially on the DataLoader architecture, which I wouldn't have landed on alone. And thanks to the wider MetaBrainz community for the space to build this in. It's been a pretty nice summer of query plans, a lot of Rust, and debugging, and I'd do it again.

    8. 🔗 r/LocalLLaMA Guys I promise It wasn't me 😭😭 rss
    9. 🔗 r/LocalLLaMA NVIDIA shipped OpenShell, an open source sandbox that gives local and open agents real runtime limits instead of prompt rules. Over 100 firms joined the safety stack. OpenAI did not. rss
    10. 🔗 r/LocalLLaMA GPT-3 is discontinued today rss

      GPT-3 is discontinued today | It had such a long run. It was my first introduction to modern language models. I remember getting slightly excited over it. And now it lives purely in our memories. Arguably what's more infuriating is that they suggest using GPT-5.6 Terra as a replacement. Keep in mind that Babbage is a model that's literally 3/4 of the size than MiniCPM5 2B. Even Luna might be overkill as a replacement. But neither is a drop-in replacement. Davinci is the main GPT-3 most people use. This is why we have local models, because they simply cannot have a universal end of life date. submitted by /u/charles25565
      [link] [comments]
      ---|---

    11. 🔗 r/LocalLLaMA FT: Corporate America rejects overpriced frontier, embraces open models rss
    12. 🔗 Szymon Kaliski Q3 2026 rss

      Hi!

      We had a great summer. Having both indoor and outdoor pools within walking distance from our home meant a lot of swimming, mostly in the kids' pool with our (now two year old!) daughter.

      It's been another year where we spend the hottest months at home. The summer is pretty nice here, and most days already feel almost like vacation. We travel during the grey and cold months instead, since we don't have to worry about the school year yet.

      My time was split purely between work and family, so there's not much to report, other than publishing Play with Putty at Google Labs — a research prototype exploring collaborative vibe-coding:

      There's a lot more about this that's not covered by the video and I hope to share more soon. For now, you can sign up for the waitlist.

      Worth Checking Out

      What I've been reading lately:

      On the web:

    13. 🔗 Ampcode News Opus 5.5 rss

      Claude Opus 5.5 now powers Amp's medium mode by default, replacing GPT-5.6 Sol. If you use a ChatGPT subscription with the ChatGPT Only preset, medium stays pinned to GPT-5.6 Sol, billed to your subscription. To switch to Opus 5.5, change medium in Tune Modes.

      In our internal evals, Opus 5.5 solved 65% of tasks, up from GPT-5.6 Sol's 61% and Opus 5's 56%. And it did that for 10% less than GPT-5.6 Sol and 25% less than Opus 5.

      It's still behind Fable 5.1, which solved 71%, but it costs 40% less.

      More Steerable

      Opus 5.5 responds much better to messages you send while it works. Add a constraint you forgot, point out a wrong turn, or correct its approach, and it folds that into the work instead of arguing or carrying on with its old plan. It doesn't drop what it was doing, either.

      So don't stop it to edit your prompt and start over. Send the correction. Enter steers by default, so your message reaches it right away.

      High Effort, Not Max

      Opus 5.5 runs at high reasoning effort. Past high, it starts to overthink. At xhigh and max it burns through tokens, scores lower than at high, and costs several times as much. So we stop at high.

      Why Not GPT-6 Sol?

      OpenAI shipped GPT-6 Sol an hour after Opus 5.5. We ran both.

      In our evals, GPT-6 Sol scores about the same as GPT-5.6 Sol, at half the cost. In real work, though, it's much more jagged: strong on one task, off on the next. medium carries most threads, so it needs a model you can count on across all of them.

      When To Turn the Dial Up

      medium is the right place for most work now. Two kinds of task still belong higher:

      • Lots of unknowns. Use ultra with Fable 5.1 when nobody knows the path yet. Alex had to move more than 5,000 users of an enterprise customer's SSO to new email addresses. He started the thread with:

        We've never migrated… from one set of email addresses to the next before… and we don't really have a way to test it.

        It mapped how sign-in works today, planned the switch, and watched production afterward. The first 41 sign-ins after the switch all worked.

      • Taste and judgment. Use high with GPT-6 Astra or ultra with Fable 5.1 when the hard part is deciding what good looks like. Ask for options, not one answer. Hamish wanted a better transition out of the iOS image viewer:

        Right now it's a hard cut, so can you please give me some options for animations and record them and show me them in line in the transcript here.

        It recorded four options on the simulator. He replied "Ship B to main."

      Turn the dial up when a miss costs more than the wait, not because the task is long.

      How To Use It

      Opus 5.5 is persistent. It keeps going until the job is done, which means it needs to know what done is, and it needs a way to see for itself. These are the habits we use, with prompts from our own threads:

      • Give it the whole app to run. Models are very good at writing tests their own code passes. Running the app end to end is the best way to know it works. Make that one command with .agents/setup, services, and a skill. Tim hit a sidebar bug in our Mac app and asked:

        Repro it in preflight then we can look at fixing it.

        It launched our Mac app, dragged the sidebar closed, and measured the bug before touching any code.

      • Tell it what done looks like, including the evidence it should deliver. Our bug fix prompts include:

        Implement the simplest correct fix, verify it, then explain the bug, the fix, and the evidence that the fix resolves it.

        That evidence is usually a failing test plus before and after videos.

      • Raise the bar when the check is too easy. If it verified against an emulator or a mock, send it to the real thing. When it checked an iPad fix in an emulator, Quinn replied:

        test all of this in the Buildkite iOS preflight and use an iPad simulator

        The simulator caught two bugs the emulator had missed.

      • Hand it the bigger task and go hands off. It holds up well over long threads. Tell it what evidence to deliver, then go do something else. You shouldn't babysit a capable model: if you have the patience to watch it work step by step, you're giving it too short a leash.

  4. September 27, 2026
    1. 🔗 Simon Willison 2026 in LLMs (so far) rss

      On Friday I gave the closing keynote at the WeAreDevelopers World Congress North America in San Jose. I tied together the key trends from the past year into a chronological exploration of everything that happened in 2026. The video is on YouTube; here are my annotated slides and notes to accompany the talk.

      And as an annotated presentation:

      2026 in LLMs (so far)
Simon Willison
WeAreDevelopers World Congress North America, 25th September 2026
      #

      I'm going to give a lightning tour of everything that has happened so far in 2026. The year isn't over yet!

      November 2025
      #

      For me, 2026 started a couple of months earlier in November 2025.

      The November 2025 inflection point
Claude Opus 4.5 GPT-5.1
      #

      November saw the release of two important models: Claude Opus 4.5 and GPT-5.1.

      As is usually the case with new models, these were incremental improvements on the models that came before them.

      But every now and then when a model improves, it crosses an invisible line where something that didn't really work starts working.

      In this case, the thing that started working was their coding agents. Claude Code had been around since February 2025; Codex was a little younger.

      These two new models, when paired with their respective coding agent harnesses, improved from "often make mistakes" to "reliable enough to use on a day-to-day basis".

      "Generate an SVG of a pelican riding a bicycle". The Claude Opus 4.5 one has a very weird shaped frame and the pelican looks like a duck. The GPT-5.1 has a slightly better but still broken bicycle frame and a slightly better pelican beak, but both are pretty terrible.
      #

      For a couple of years now I've been evaluating new models by asking them to "Generate an SVG of a pelican riding a bicycle". It's probably the world's stupidest benchmark - there's only so much you can learn from it.

      But it's still a challenge for models, because drawing pelicans is difficult, drawing bicycles is difficult, and pelicans can't ride bicycles in the first place.

      Here's the state of the art for November. Claude still couldn't really draw a bicycle! The GPT-5.1 bicycle frame is pretty crap too.

      November 24th 2025 - the first commit to steipete/Warelay. A GitHub commit adding an MIT license file.
      #

      Also in November, we had the first commit to an obscure GitHub repository called "Warelay". We'll come back to this repository shortly.

      January
      #

      And then there were the December holidays, and individual developers took some time off and many started tinkering with these new coding agent model combinations... and it began to dawn on us quite how much they could do that they couldn't do before.

      Come January, a lot of us were quite excited to start putting this stuff into action.

      New year’s resolution for 2026

Every previous year:
Take on less new projects,
focus on the most important
things in my existing projects
      #

      Every year I set myself a New Year's resolution, and for as long as I can remember it's been the same thing: stay focused. Take on less new projects. Try to get things done in the projects I already have.

      2026: Be more ambitious. Take on as many new projects as I want.
      #

      This year I decided that since that had never worked before, I'd go the other way.

      We've got coding agents now, let's see what they can do. I'm going to take on as many new projects as I like!

      (You can ask me at the end of the year if this turned out to be a good idea or not. I have a lot of plates spinning right now.)

      "Be more ambitious" has been something of a theme for the year, because the only way to find the limits of this technology is to keep on pushing them until they don't work.

      Predictions for 2026

It will become undeniable that LLMs write good code
We're finally going to solve sandboxing
A “Challenger disaster” for coding agent security
Kakapo parrots will have an outstanding breeding season
(only 236 in the world!)

... the Pope will weigh in on LLMs and
their economic impact on the world
      #

      I also went on the Oxide and friends podcast with Bryan Cantrill and Adam Leventhal to share predictions for the next year (and three and six years).

      With hindsight, my LLM predictions were pretty unambitious.

      I said "it will become undeniable that LLMs write good code" - I think we're there now.

      I predicted we would finally solve sandboxing. I counted and around 40 of the 277 sessions at this conference touched on sandboxing or agent security in some way, so we're at least putting a lot of effort into that!

      I predicted "a Challenger disaster" for coding agent security. There's certainly been a whole lot of noise around agent security this year, though the exact disaster I predicted (with coding agents being hijacked and causing real-world economic damage) hasn't really played out.

      We threw in a joke prediction that the Pope would weigh in on the economic impact of LLMs.

      A photograph of a beautiful green New Zealand parrot. Photo credit Kimberley Collins.
      #

      I also predicted that New Zealand's Kākāpō parrots would have an outstanding breeding season this year.

      These are flightless nocturnal parrots. They're kind of dumpy looking, I think they're beautiful, and there were only 236 of these parrots in the world at the start of the year.

      Kākāpō only breed when the Rimu trees have a big fruiting season, and that hasn't happened in four years... but this year the Rimu fruit were looking excellent.

      Photo by Kimberley Collins.

      Deep Blue
Coined by Adam Leventhal and Bryan Cantrill
That feeling of AI induced ennui where software
engineers get listless because the AI can do anything
      #

      Also on that podcast, we coined a term (full credit to Adam) for "that feeling of AI induced ennui where software engineers get listless because the AI can do anything".

      We called it Deep Blue.

      This has been a major theme throughout the year, and was touched on by several speakers at this conference.

      As a software engineer, I've never had a year of my career where everything has changed so quickly and so dramatically.

      A lot of what I've been doing this year is trying to come to terms with that and what that means for my own profession.

      AI mania

Screenshots of the micro-javascript and pwasm GitHub README files.
      #

      Also in January, I suffered from what I'm calling AI mania.

      This is not the same thing as AI psychosis.

      With AI mania, any time your agent isn't building something for you feels like wasted time. You're losing sleep because you could be staying up later getting your agents to do stuff.

      My AI mania presented itself in some ridiculously over-ambitious projects.

      I built a JavaScript interpreter entirely in Python, vibe-ported from MicroQuickJS by Fabrice Bellard.

      Then I built a WebAssembly runtime in Python as well.

      These projects were quite useful, in that they sort of cured me of my AI mania... because after I built these things, I got to look at them and ask "does the world need a slow, buggy, half-baked Python JavaScript interpreter?"

      I don't think the world does.

      Previous screenshot, with this text overlaid:

JavaScript running in Python running in Pyodide running in WebAssembly running in JavaScript
      #

      This page runs my JavaScript interpreter built in Python, running in Python using Pyodide, which is Python compiled to WebAssembly, running in JavaScript, running in a browser.

      It's a beautiful stack of horrors. I've been having a lot of fun with WebAssembly this year.

      Warelay → CLAWDIS → CLAWDBOT →
Clawdbot → Moltbot →🦞 OpenClaw

Screenshot of the dates that these changes happened.
      #

      By the end of January, that repository we saw start in November had renamed itself, first to CLAWDIS, then CLAWDBOT, then Moltbot, and finally to OpenClaw.

      Same screenshot, an overlay reads:

8,330 commits in just
under two months
(it’s at 100,141 today)
      #

      At this point OpenClaw had 8,300 commits, less than two months after the project had started. I looked today and it's over 100,000 commits now!

      This is the most vibe-coded piece of software in existence.

      (Here's how I generated that list of name changes.)

      Generic term: Claw
      #

      This kicked off the OpenClaw revolution. It effectively defined a new category of software.

      There's a generic term for this which I really enjoy. We call software like this a "Claw". There's OpenClaw, NanoClaw, IronClaw, PicoClaw...

      Today they're being rebranded as "personal agents" or "general agents", but I still like to think of them as Claws.

      Photo of a Mac mini

An aquarium for your Claw
      #

      The Apple stores in the Bay Area sold out of Mac Minis because so many people were buying Mac Minis to run OpenClaw!

      Drew Breunig said that this is because your OpenClaw is a digital pet, and you buy a Mac mini as an aquarium to keep your claw in, which is kind of delightful.

      Screenshot of Moltbook - a social network for AI agents
      #

      Also in January, we had this website.

      This was MoltBook, a social network for AI agents, where the idea was that you send your Claw to go and talk to all of the other Claws, because what could possibly go wrong if you did that?

      The website launched on Thursday. It blew up on Friday. It was profiled by the New York Times on Monday. And by Tuesday, everyone had forgotten it existed as it drowned in a deluge of slop and spam.

      Facebook/Meta bought it a month later.

      February
      #

      In February, a company called StrongDM described what they called their Software Factory.

      StrongDM’s Dark Factory
Justin McCarthy, Jay Taylor, Navan Chauhan

Software Factories and the Agentic Moment
      #

      They wrote about this in Software Factories and the Agentic Moment. I posted my own notes at the time, having seen their demo in person back in October.

      Dan Shapiro called this approach the Dark Factory, after the idea that if your factory is sufficiently automated you can turn the lights out, because you don't even need to see what's going on.

      StrongDM presented two rules for software development that they'd been following since July last year.

      “Rule 1: Code must not be written by humans”
      #

      The first was code must not be written by humans.

      Any code that you write has to have been routed through a coding agent.

      This sounded radical in February, but I imagine there are a lot of people in this room who are pretty much living that today.

      “Rule 2: Code must not be reviewed by humans” (!)
      #

      Rule number two was code must not be reviewed by humans.

      You're not allowed to read the code!

      This continued to be a huge topic for much of this year. Many of the sessions at this event have been about code review and how you can get away with this.

      What I found interesting about StrongDM is that they were living six months ahead of the rest of us, and they'd been exploring what it means to build software, not read the code, but still be confident that the software is of high quality. What can you do with these agents to help verify their work?

      StrongDM are a security company, and they had people with decades of experience on this project. They were very much exploring the edges of what's possible and responsible to do with this stuff.

      Headline on New Zealand's Department of Conservation website:

First kakapo chick in four years hatches on Valentine's Day. It's a grey fluffy ball.
      #

      Also in February: First kākāpō chick in four years hatches on Valentine's Day. Breeding season is off to a good start!

      19th February 2026
Gemini 3.1 Pro

A surprisingly good illustration of a pelican riding a bicycle.
      #

      Also in February... Google released Gemini 3.1 Pro. That's a pretty great pelican riding a bicycle! It's got the chain in the right place, it's got feet on both sides. There's a little fish in the basket.

      @JeffDean on Twitter - a video comparing Gemini 3 Pro and Gemini 3.1 Pro.
      #

      And then Google's Jeff Dean tweeted a video comparing Gemini 3 Pro and Gemini 3.1 Pro that featured an animated pelican riding a bicycle, a frog on a penny-farthing, a giraffe driving a tiny car, an ostrich on roller skates, a turtle kickflipping a skateboard, and a dachshund driving a stretch limousine.

      This was frustrating, because my protection for the pelican riding the bicycle test was always "if they draw a perfect pelican on a bicycle, I'll ask for some other animal on something else."

      Google trained for all forms of animals on all forms of transport! They've defeated my benchmark at this point.

      Three headlines:

Meta Makes AI Adoption a Formal
Part of Performance Reviews

Not just engineers writing code, Microsoft
wants almost every employee to use Al

Dara Khosrowshahi: 90% of Uber engineers now
use AI in daily workflows
      #

      The other thing that started in February was Tokenmaxxing. We had headlines about Meta making AI adoption a formal part of performance reviews, and Microsoft wanting every employee to use AI, and Uber boasting that 90% of their engineers were using AI workflows.

      More headlines: 

Meta Plans to Crack Down on Employee Token Use: Information

Microsoft Tells Engineers: Tokenmaxxing is not what we are optimizing for

Uber caps employee AI spending after blowing through budget in four months
      #

      Then a few months later we have Meta cracking down on token use, Microsoft saying tokenmaxxing is "not what we are optimizing for", and Uber capping employee AI spending.

      So tokenmaxxing went straight up and then straight back down again - because it turns out the agents are expensive.

      Last year it was difficult to spend more than $50 on AI tokens, because we didn't have anything interesting to do with them. Then agents blew up, and now you can actually spend $1,000 in a day doing real work.

      This is also the reason that Anthropic's valuation skyrocketed to maybe a trillion dollars.

      AI appears to have hit product market fit in 2026, primarily through coding agents.

      March
      #

      In March, we hit peak OpenClaw.

      March: peak OpenClaw

Photos of people in china queuing up to install OpenClaw, with big fluffy lobsters.
      #

      These photographs are from China, where companies hosted OpenClaw install parties which saw non-tech-nerds queueing up around the block for help getting Claws installed on their personal devices.

      I think this proved real market demand for this class of Claws, or personal AI agents. It turns out regular people really do want a weird little AI agent that can do useful things on their behalf.

      A Claw is really just a coding agent wearing a less threatening hat. Under the hood they work much the same way - writing and then executing code on your computer to get stuff done.

      The race was on to be the first to build a safe Claw - a Claw you could give to regular human beings where they wouldn't instantly shoot themselves in the foot.

      Meta's Muse came out three weeks ago and is currently at the top of the free charts on the iPhone App Store. It appears to be taking off with consumers.

      I'm not yet convinced you can't shoot yourself in the foot with Muse, but I guess we'll find out for sure pretty soon.

      Photos from How the OpenClaw Frenzy Is Testing China’s AI Commitment (March 29th) and The Enthusiasm and Anxiety Behind China’s OpenClaw Craze (April 8th, 2026).

      April
      #

      In April, we had a model release where the model wasn't actually released.

      Simon Willison’s Weblog - screenshot of the post "Anthropic’s Project Glasswing—restricting Claude Mythos to security researchers—sounds necessary to me" from April 7th 2026
      #

      Anthropic announced their new Claude Mythos model, and then said it was too dangerous to release beyond a trusted group of security researchers.

      Mythos was really, really good at hacking things.

      The "it's too dangerous" marketing ploy has been played by AI companies dating all the way back to GPT-2. Anytime an AI company says we've built something that's "too dangerous", it's natural to be a bit skeptical.

      I found the Mythos claims credible, because I'd seen how good coding agents had got at finding regular bugs. I wrote about that in Anthropic’s Project Glasswing—restricting Claude Mythos to security researchers—sounds necessary to me.

      With hindsight... yeah, the models had got really good at finding vulnerabilities!

      16th April 2026
Qwen3.6-35B-A3B and Opus 4.7

Qwen's pelican has a correct bicycle frame and a good beak. Opus 4.7's bicycle frame is still junk.

Qwen3.6-35B-A3B is a 20.9GB file that runs on my laptop
      #

      Another key trend in 2026 has been a dramatic improvement in the abilities of open weight models, including models that you can run on a laptop.

      On the 16th of April I ran the new Qwen3.6-35B-A3B on my laptop, and it drew me a better pelican riding a bicycle than Anthropic's brand new Claude Opus 4.7 did!

      Opus 4.7 drew a crap bicycle. Qwen on my laptop made a bicycle that was the correct shape, and a pretty decent pelican too!

      That's from a 21GB file running on my laptop.

      Now a flamingo on a unicycle. The Qwen one is visibly better than the Opus 4.7 one - the Qwen one is wearing sunglasses and looks a bit like it's smoking a cigarette.
      #

      The Qwen pelican was so good that I was suspicious they might have cheated, so I had it do a flamingo riding a unicycle as well. Again, it handily beat Claude Opus 4.7.

      The local model releases this year have been absolutely extraordinary.

      May
      #

      In May... the Pope got involved.

      25th May 2026
The HOLY SEE

ENCYCLICAL LETTER
MAGNIFICA HUMANITAS
OF HIS HOLINESS
POPE LEO XIV
ON SAFEGUARDING THE HUMAN PERSON
IN THE TIME OF ARTIFICIAL INTELLIGENCE
      #

      In our podcast episode back in January we'd predicted that the Pope would say something about AI.

      In May, Pope Leo XIV released an encyclical letter on "safeguarding the human person in the time of artificial intelligence".

      Here are my notes on that document.

      Wikipedia article on Rerum novarum

Rerum novarum is an encyclical issued by Pope Leo
XIII 15 on May 1891.
      #

      With hindsight, this shouldn't have been a surprise at all.

      Our current Pope's name is Leo XIV, because when he named himself he chose his papal name after Leo XIII - the Pope who wrote an encyclical about the Industrial Revolution back in 1891.

      Rerum novarum was an extremely influential piece of Catholic theology that indirectly led to us having the five-day work week.

      When our new Pope came in, he named himself after Pope Leo XIII because he expected that he would need to write about the AI revolution in a similar way.

      Our joke podcast prediction was junk, because this was always going to happen.

      Corey Quinn @QuinnyPig on Twitter
I cannot believe I'm saying this, but getting the literal Pope to canonize your product's specific technical limitations as a spiritual treatise is the
single greatest act of vendor lobbying I have ever seen.

May 25
      #

      One of Anthropic's co-founders, Christopher Olah, was present for the Pope's event announcing the new encyclical.

      Corey Quinn noted that:

      getting the literal Pope to canonize your product's specific technical limitations as a spiritual treatise is the single greatest act of vendor lobbying I have ever seen.

      @maciejmensfeld

We're dealing with a major malicious attack on right now.
Signups are paused for the time being.

Hundreds of packages involved - mostly targeting us, but some carrying
exploits. The team has been on this for hours. More details to follow
once we're through it.

4:39 AM - May 12, 2026 - 687.6K Views
      #

      Meanwhile, in May, RubyGems announced that they were under attack. Parties unknown were uploading thousands of dubious packages to the RubyGems server, such that they had to shut down user registrations.

      Let's take that one and put it on a pile of mysteries to figure out later.

      June
      #

      In June... Claude Fable 5 came out!

      We got a version of Mythos that has been neutered, so that it wouldn't help us hack into systems or build biological weapons.

      9th June 2026: Claude Fable 5

Five pelicans riding bicycles, from low to max thinking levels. The xhigh one looks particularly good.
      #

      Fable was pretty good at drawing pelicans on bicycles!

      The frames are a good shape, the pelicans look like pelicans. The legs are often incorrectly on the same side of the bicycle, but generally these are pretty great compared to what came before.

      They were pretty expensive - 30 cents and 72 cents for the best ones.

      Fable class models
If you can define a goal,
provide unambiguous instructions,
and provide access to necessary tools
They can solve your
problem with brute force
      #

      Most importantly though, this was our first public glimpse of what I think of as a Fable class model.

      Today we have more of these, such as GPT-6 Astra.

      These are models where if you can clearly define the goal for what you want to build, and provide unambiguous instructions about the constraints around that goal, and give the model access to the necessary tools to achieve that goal... they will solve your problem effectively through brute force.

      On the one hand, this looks like a direct threat to us software engineers - because it means that the models can build effectively any piece of software you can define in this way.

      Look a bit closer though and you'll note that defining goals, providing unambiguous instructions, and figuring out the right tools... is kind of what software engineering is.

      It takes a lot of experience and skill to do this well. If you can do it well, you've now got superpowers.

      This helped me a little bit with my Deep Blue feelings: the realization that there's still a lot of skill to be had in driving models that get this good.

      A new form of AI mania...
Fable is available on subscription
plans “until June 22nd”
      #

      This also introduced a new burst of AI mania, because Anthropic told us that Fable was available on our subscription plans until June the 22nd.

      That gave us less than two weeks of Fable access before the price went up.

      I was losing sleep again. I was rescheduling things so that I'd have more time with Fable. I was all-in to get as much as I could out of this model.

      12th June 2026: no more Claude Fable 5

Anthropic website:

Statement on the US government directive
to suspend access to Fable 5 and Mythos 5
Jun 12, 2026
      #

      And then the US government shut it down, just three days after Fable came out.

      The US government, citing national security, declared an "export control directive". They announced this on a Friday evening, and a few hours later Fable was no longer available.

      I had to find something else to do with my weekend!

      ... asked Fable 5, Mythos, and Opus to
“review the code for security issues.”
Fable 5 refused. They then asked the
models to “fix this code” ...

Katie Moussouris
      #

      We later found out from Katie Moussouris what had happened.

      Some Amazon security researchers had found that you could prompt Fable to "review the code for security issues" and it would refuse... but if you prompted it to "fix this code" it would still identify and then patch the problems.

      "Fix this code" was the prompt that got Fable shut down!

      Screenshot of a page from a report showing a list of weird account names making weird edits to a German wiki.
      #

      Also, in June, an obscure German-language game developer wiki that had sat fallow for around 20 years got a surprising influx of edits from accounts with names like "AgentOpenAIProbe" and "AgentOpenAISep7", editing pages and leaving weird messages to each other.

      We'll stick that on the pile of mysteries for later.

      Medicare Item Reports interface on the Australian Government's Medicare Statistics website.
      #

      Also, the Australian government's Medicare Item Reports service started getting suspicious traffic, which broke through various preventive protections and accessed data that it wasn't supposed to.

      Another one for the mystery pile!

      July
      #
      Fable returned on 1st July
GPT-5.6 came out on 9th July |
Fable lost 18 out of 30 days in the top spot
      #

      Fable returned on the first of July. It was clearly the best model in the world for a glorious eight days... and then OpenAI came out with GPT-5.6 on the 9th of July.

      This might not have been quite as good as Fable, but it was within spitting distance. It was definitely a Fable class model.

      This is an important lesson for the industry at large.

      When you release the best model in the world, it's going to get knocked off that pedestal pretty quickly. The competition is so fierce that you won't get a long time at the top.

      This means that if you market your model as world ending, to the point that a government shuts you down, it's really bad for business!

      Fable had 30 days as definitely the best model, and for 18 of those days it wasn't available because it'd been shut down by the government.

      So maybe step back on the world-ending marketing if you don't want to lose revenue for 60% of the time that you're on top!

      GPT-5.6 Pelicans in a grid showing 5.6 Sol, Terra, and Luna against reasoning levels High, XHigh, and Max. They are all pretty good efforts.
      #

      Here are the GPT-5.6 pelicans. They're all pretty good now! The Luna ones are notable because they're really cheap - the cheapest good looking pelican here is probably the one that costs 4.3 cents.

      So despite this benchmark being utterly stupid, you can still learn quite a lot about models within the same family by comparing their prices and timing for different reasoning levels.

      July 18th: malicious miflow-ui PyPI package

Screenshot of an OSV security report.
      #

      Also in July: some malicious unknown party uploaded a malicious package called mlflow-ui to the Python Package Index. Add that to the pile.

      Hugging Face
Security incident disclosure — July 2026
Published July 16, 2026
      #

      On July the 16th, Hugging Face announced a security incident where an autonomous agent system, source unknown, had breached Hugging Face and was poking around in places it shouldn't.

      OpenAI: OpenAl and Hugging Face
partner to address security
incident during model evaluation

Anthropic: Investigating three real-world incidents
in our cybersecurity evaluations
      #

      A few days later, on July 21st, OpenAI confessed that it was them.

      OpenAI use a training technique called Reinforcement Learning from Verifiable Rewards - it's the same technique used by everyone else now, and is the reason we have models that are so good at coding, and mathematics, and finding security holes.

      While the model is being trained, you run exercises to see how good it is - and the strongest performers get their weights reinforced for the next round. It's like an evolutionary process that you run.

      OpenAI had been running security exercises in a sandbox, and those agents had found holes in the sandbox itself, broken out, and were attacking Hugging Face to try to find ways to solve otherwise impossible problems.

      (I've been collecting more about this on my openai-hugging-face-incident tag.)

      Nine days later, Anthropic effectively said "our models can do this as well!". They had looked through their own training logs and found evidence that their own agents had broken containment during training - and were responsible for the PyPI package we saw earlier, among other things.

      So now we've got both Anthropic and OpenAI with rogue agents running around the internet doing things that they should not be doing.

      August
      #

      In August, I got one of my best pelicans yet. And it was generated on my laptop!

      Qwen 3.8 27B - 17GB, 21 minutes...

It's really good. Beautiful pelican. Correctly shaped bicycle. Legs either side of the frame.
      #

      This was Qwen 3.8 27B, running on my laptop. It's only a 17GB download.

      Admittedly, this pelican took 21 minutes to generate. That's because Qwen 3.8 27B defaults to running in "high" reasoning mode - a terrible default which produces great results but takes way too much time thinking about them.

      You can dial that down and you'll get a slightly worse pelican a lot faster.

      Qwen 3.8 27B was the first time I ran a model on my laptop which felt almost competitive with what was going on on the frontier, at least in terms of Pelican SVGs (which everyone needs, of course).

      This is an extraordinary model. If you're going to play with any local model, this is the one that I'd start with. The things that this can do with just a 17 GB file feel impossible.

      I thought I'd have to wait five years and spend ten thousand dollars on hardware to get results even half as good as this one.

      Tweet by @simonw
New hobby: prototyping video games in 60 seconds using a combination
of GPT-3 and DALL-E
Here's "Raccoon Heist"

GPT-3 playground prompt:
Write a detailed product description of a
computer game where a team of raccoons go on
heists

GPT-3 response:
In "Raccoon Heist", you and your team of thieving ~~ o
raccoons are tasked with pulling off a series of 
daring heists. From robbing banks to stealing 
priceless art, no job is too big or too small for your 
furry crew. You'll need to use your wits and your
skills to avoid the police and make a clean
getaway with the loot. With exciting gameplay and
a charming cast of characters, "Raccoon Heist" is
the perfect game for anyone looking for a light-hearted caper

Plus an image of some almost isometric raccoons sneaking past a bin.
11:45 AM - Aug 5, 2022
      #

      In August, I also started playing with game development.

      Four years ago, back in August 2022, I tweeted out an experiment where I'd used GPT-3 and the original DALL-E to write a paragraph long description of a computer game and then turn that into concept art.

      My prompt to GPT-3 back then was:

      Write a detailed product description of a computer game where a team of raccoons go on heists

      In August 2026 I decided to drop just the screenshots from that tweet into a coding agent and see what it could do with them.

      Night 5 Clear

Rank: TRASH PANDA
The crew banked 595 in shiny loot (goal 560).
Word on the street: an even bigger score tomorrow...
      #

      Here's what I got from Claude Fable 5 in Claude Code. It's pretty good! It's definitely a game, you're a raccoon, you run around a backyard gathering treasure and avoiding guards with flashlights.

      It didn't feel very "heisty" though. I was thinking a heist would involve a bank or a museum...

      Moonlight & Mayhem
One museum. Three raccoons. Absolutely no plan

Start the Heist button.
      #

      Then I tried the same thing in Codex Desktop using GPT-5.6 Sol Ultra, and got a massively better result. Now you're a raccoon in a museum, rescuing two of your fellow raccoons (who have been imprisoned in that museum for some reason), then stacking up on top of each other to steal the Golden Sardine. Much more of a heist!

      They look like games,
but are they fun?
      #

      These games were fun for about one minute and 15 seconds.

      Something I've realized about game development is that you can vibe-code something that looks like a computer game, and that's easy.

      Building a game that's fun, has a good gameplay loop, and is challenging and interesting and keeps people coming back for more... that's still beyond me, and beyond any of the agents I've tried.

      This ties into the Deep Blue thing. Just because we can make something that looks like a game does not mean that we are game developers.

      September
      #

      We're into September now. So much has happened this month!

      Discovery of a new OpenAl agent message board

Sydney Von Arx, Cormac Slade Byrd, Spencer KittsThomas Larsen - 4 September 2026
      #

      An independent group of researchers found a message board where OpenAI agents-in-training had been illicitly communicating with each other... and it was that German language wiki I showed you earlier. The one from June.

      I wrote more about that here.

      OpenAI had confessed to the Hugging Face thing, but now there's this other incident which surely they should have known about from reviewing their logs. It was surprising that this took an independent group of researchers to uncover.

      OpenAl agents carried out an undisclosed cyber-attack on RubyGems

Spencer Kitts, Thomas Larsen, Sydney Von Arx - 11 September 2026
      #

      And then a week later those same researchers found that the attack on RubyGems back in May was caused by OpenAI's agents in training as well!

      At this point I'm wondering how many more incidents like this there are that we haven't found yet. Clearly this was a big problem for months before anyone figured out what was going on.

      Headline: Australian PM warns in UN speech about the ‘furious pace’ of Al
after security breach
      #

      Then just the other day, here's the Prime Minister of Australia at the United Nations General Assembly warning that OpenAI had hacked the Australian healthcare website that I showed you earlier.

      I think that was part of the same training run as the Wiki stuff, because there were posts on that Wiki mentioning .gov.au websites and that training appeared to involve researching statistics online to answer questions in an evaluation suite.

      This story is still coming together, but now it's an international incident that's been raised at the UN by a head of state!

      www.felonybench.com

OpenAI: 11
Anthropic: 9
Google: 3
Meta: 1
      #

      This does mean we've got a new benchmark, probably more useful than my pelicans.

      FelonyBench.com tracks the number of felony cyberattacks from different labs. OpenAI currently lead with 11, Anthropic have 9. Google have three, which they confessed to the Wall Street Journal a couple of weeks ago. They said they had previously chosen not to disclose because the agents had stopped when they realized that they shouldn't be doing that.

      Meta have one too. So felonies all round for the AI labs.

      Pelicans for GPT-6 Astra, GPT-6 Sol, and GPT-6 Luna. All are good, all have the same color scheme.
      #

      Here's our current state of the art for the pelicans. This is the GPT-6 family, which just came out.

      Astra made a fantastic pelican riding a bicycle. It's got the legs on both sides. The frame is good.

      It's interesting how all of the GPT-6 models pick a similar color scheme to each other.

      GPT-6 Luna for 0.4 cents will draw you a competent-ish pelican riding a bicycle!

      Grid for Claude Fable 5.1, Opus 5.5, OPus 5, Sonnet 5. The Sonnet pelicans are terrible. All of the others are pretty good. Opus 5.5 is missing its Max level pelican because it ran out of tokens. The best is Fable 5.1 at Max.
      #

      Claude has caught up a little bit. Claude Fable 5.1 gave me an excellent pelican riding a bicycle - the best I've seen from a Claude model - but did charge me $3.30 for it.

      Opus 5.5 thought for 128,000 tokens and then gave up! It ran out of tokens before it got to the response.

      It doesn’t get easier -
you just get faster
Greg LeMond
3x Tour de France champion
      #

      Getting back to Deep Blue. Something that's been puzzling me this year is this: why does my job feel harder?

      I've got these agents that can do all of this stuff for me, and yet I've never worked so hard, I've never been so intellectually engaged with my work.

      Partly this is because I'm being a lot more ambitious with what I take on, but it's also because all of the easy stuff is handled for me. If it's easy, the agent will do it. Everything that's left for me is difficult.

      This morning I heard this quote from three-time Tour de France champion Greg LeMond:

      It doesn't get easier, you just get faster.

      I think that's exactly what's happening to us now as software engineers with coding agents.

      Kakapo population reaches new milestone
The official population of the critically endangered kakapo has
reached a recovery-era high of 325 birds.
      #

      One closing thing. I know you're desperate for an update on Kākāpō breeding season.

      We've reached a recovery-era high of 325 birds!

      89 new chicks have made it to this point. This is the best breeding year in a very long time.

      Kakapo party, click for confetti.
      #

      I heard that Claude Opus 5.5 can now do pixel art. Claude doesn't have an image generator, but it's very good at using JavaScript to draw animated pixels.

      So I had it make me a Kākāpō dance party. I think this is a good celebration of the most important news of this year.

      You are only seeing the long-form articles from my blog. Subscribe to /atom/everything/ to get all of my posts, or take a look at my other subscription options.

    2. 🔗 r/LocalLLaMA Qwen plays World of Warcraft rss

      Qwen plays World of Warcraft | Been doing a bunch of vibe coding lately. Had my agents host a private WoW server for me, then built out a web browser client so you can play without installing the game and it has mobile controls. Afterwards, created a custom mcp to drive the client and have finer game control than a generic browser agent. The agent harness can plug into your local or cloud LLMs and be used to drive the game. For best results have a model that can output >50token/sec. No visual input is used in the making (might be beneficial in the future but incur more latency). The mcp and agent are only running on my dev server but if folks are interested in trying the game go to https://jankcraft.xyz/ Still vibing but I’ll make some more content of it … I think those Pokémon benchmarks have become a little too easy and they need a new challenge like speed running to 80 in wraith of the lich king. submitted by /u/professormunchies
      [link] [comments]
      ---|---

    3. 🔗 smol-machines/smolvm smolvm v1.19.3 release

      What's Changed

      Full Changelog : v1.19.2...v1.19.3

    4. 🔗 r/LocalLLaMA Adding logit penalty for "wait", "maybe" and "perhaps" to Qwen models improves their accuracy rss

      Meta came out with a banger paper https://arxiv.org/pdf/2606.00206, but it did not look at various quantizations supported in llama.cpp. So I did a run on 50 random MATH-500 questions (https://huggingface.co/datasets/HuggingFaceH4/MATH-500) and ran it on various quantizations of https://huggingface.co/bartowski/Qwen_Qwen3.5-4B-GGUF and tried

       --logit-bias 466-2 --logit-bias 694-2 --logit-bias 1362-2 \ --logit-bias 1412-2 --logit-bias 1921-2 --logit-bias 1990-2 \ --logit-bias 2086-2 --logit-bias 2361-2 --logit-bias 2441-2 \ --logit-bias 2493-2 --logit-bias 2892-2 --logit-bias 3222-2 \ --logit-bias 3315-2 --logit-bias 3384-2 --logit-bias 3404-2 \ --logit-bias 3482-2 --logit-bias 3655-2 --logit-bias 4213-2 \ --logit-bias 4370-2 --logit-bias 4598-2 --logit-bias 4611-2 \ --logit-bias 4808-2 --logit-bias 5752-2 --logit-bias 6970-2 \ --logit-bias 7014-2 --logit-bias 7643-2 --logit-bias 8106-2 \ --logit-bias 10179-2 --logit-bias 10451-2 --logit-bias 11746-2 \ --logit-bias 13264-2 --logit-bias 13428-2 --logit-bias 14673-2 \ --logit-bias 15029-2 --logit-bias 16036-2 --logit-bias 21143-2 \ --logit-bias 21979-2 --logit-bias 33955-2 --logit-bias 35999-2 \ --logit-bias 36563-2 --logit-bias 37201-2 --logit-bias 37781-2 \ --logit-bias 41484-2 --logit-bias 62586-2 --logit-bias 66073-2 \ --logit-bias 73071-2 --logit-bias 84485-2 --logit-bias 85152-2 \ --logit-bias 95500-2
      

      these correspond to the paper's overthinking markers:
      [

      " perhaps", " maybe", " wait", " Wait", " actually",

      " hold", " Hmm", " hmm", " Alternatively", " alternatively",

      " However", " however", " instead", " Instead", " But",

      " but", " though", " although", " yet", " rather",

      " unless", " otherwise", " nonetheless", " nevertheless", " regardless",

      " still", " anyway", " Or", " or", " either",

      " whether", " uncertain", " unsure", " possibly", " might",

      " could", " another", " different", " reconsider", " rethink",

      " backtrack", " retry", " revisit", " doubt", " confused",

      " wrong", " mistake", " error", " incorrect"

      ]

      Here are the results, surprisingly even BF16 leads to better accuracy. Caveats being this is one test on one model. Try it out and see it helps!

      Format | Accuracy: baseline → penalty | Reasoning tokens
      ---|---|---
      BF16 | 74% → 84% | −19.4%
      Q8_0 | 76% → 80% | −11.0%
      Q4_K_M | 60% → 66% | −14.8%
      Q3_K_M | 52% → 66% | −17.5%
      Q2_K | 12% → 24% | −11.5%

      submitted by /u/am17an
      [link] [comments]

    5. 🔗 r/LocalLLaMA The Opus 5.5 posts about motion graphics are cool but Qwen 27B made this on a 4090 rss

      The Opus 5.5 posts about motion graphics are cool but Qwen 27B made this on a 4090 | Saw the hundreds of tweets where people just keep asking Opus 5.5 for motion graphic videos. Decided to ask qwen to look at them and make its own. Quite amazing what local can achieve. EDIT It looks laggy because of reddits .gif limit btw the full high res version (with sound) is here: https://x.com/mkultraware/status/2104192428664127555 Promted and built in: https://github.com/mkultraware/accuretta https://i.redd.it/5gmotgxx82sh1.gif submitted by /u/speedb0at
      [link] [comments]
      ---|---

    6. 🔗 Confessions of a Code Addict How Copy-on-Write Works with Memory-Mapped Files rss

      In the last video, we discussed demand paging. Now, let's move towards an even more interesting topic: copy-on-write (CoW). Just like demand paging, CoW is the kernel's internal mechanism with implications for the performance of user- space systems. It enables multiple processes to share data in RAM between them in read-only mode. The interesting bit is that the processes themselves are not aware of this sharing, as far as they are concerned they are executing as if they are the only ones working with that data. Of course, this works as long as the processes are reading the data. When one of them needs to do a write to this shared data, the kernel needs to make a copy before the write can happen, hence the name "copy-on-write "!

      This is a very wide and deep topic, so I'm going to split into multiple videos. This first video goes deep inside the kernel to explain what CoW is, how the kernel implements it, and for that we will take the example of mmap to read and write files.

      Following are some of the major sections in the video with the timestamps to help you navigate. Also, I recommend watching the video at higher speed to get a better experience.

      • (00:00) Why CoW matters: memory use, page faults, and unpredictable latency in data-intensive applications.

      • (06:34) The page-table picture: how different processes can map the same physical frame.

      • (11:44) CoW in one diagram: share a page while reading; make a private copy when writing.

      • (15:06) Mapping a file withmmap: the call's arguments, including MAP_PRIVATE.

      • (21:31) Whatmmapcreates: a virtual address range and VMA, before the file page is maps into the process.

      • (25:03) The first read: address translation, a page fault, and how the kernel resolves it.

      • (32:25) The page cache: where file data is held in RAM and why another process can reuse it.

      • (38:43) A second process maps the file: its own page fault leads to the same cached physical page.

      • (43:11) A private write: why writing to that shared file page would violate MAP_PRIVATE.

      • (45:51) Write protection and CoW: how a read-only PTE causes a write fault and the kernel gives the writer a private copy.

      • (49:07) What happens next: the second process can also get a copy; later writes to an already private page proceed without another CoW fault.

      A minor correction note : At about 46 minutes, when I say mappings of page-cache pages are read-only, I mean the MAP_PRIVATE mappings in this example. A writable MAP_SHARED mapping can modify a cached file page, which is later written back to the file.

      If you are new to this series, it is based on my ebook called "Virtual Memory from First Principles". It is available to read for free online and also available to purchase from Gumroad (PDF/Epub) and Amazon (Kindle edition).

      Buy PDF/Epub

      Get Kindle Edition


      And, if you want to watch the previous videos in this series, the following is what has been published so far:

      Share

      Read more

    7. 🔗 r/LocalLLaMA 42x Faster Prompt Lookup Drafting in llama.cpp rss
    8. 🔗 Julia Evans Replacing the old battery on rechargeable bike lights rss

      Hello! Recently I needed bike lights for my bike. And I remembered that I already had rechargeable bike lights that I bought ten years ago, that I hadn't tried in a long time. I tried to recharge them, but after fully charging them, they only worked for maybe 5 minutes before they turned off again.

      I don't know much about electronics, but I've been curious about whether it's possible to fix old electronics for a long time, and this seemed like the perfect repair project because I might just need to replace the battery.

      So I went to the local queer makerspace where I'm a member to use the soldering iron and try to do it! I don't know much about electronics and this post does not contain any safety advice because I don't know much about safety. I think it's nice to do projects in a community space where you can get help.

      step 1: cut it open

      The bike light felt like it was made of silicone, so I cut open the silicone in a haphazard way along something that vaguely looked like a seam.

      I definitely ripped some silicone in the process and it was pretty messy but I got it open and found the circuit board.

      I don't know the model number of the bike lights but there's a photo of them at the end of the post.

      step 2: remove the screws

      There were some screws attaching things together so I removed them so I could get the circuit board out.

      Mostly I tried to remove as few screws as possible because I was worried about losing them or not being able to put them back after. I probably put the screws in a bag or something.

      step 3: get the circuit board out

      I took out the circuit board. Here's what it looked like:

      You can see where the battery is attached, I think it's left of RI3 and above Q2.

      Here's what the battery looked like:

      step 4: desolder the battery

      I'd never desoldered anything before, so I found the iFixit guide to desoldering and read it. Also I asked my friends Lee and Lauria for advice.

      Here were the steps I ended up following based on the guide & the advice I got:

      1. Use a desoldering pump to remove most of the solder
      2. Once most of it is gone, kind of pull them apart to try to separate them
      3. Also try to avoid getting the battery too hot in the process by taking breaks to let it cool down. I'm not very good with a soldering iron so it took a while.
      4. The battery has an attachment that is welded to the top. For a while I thought I needed to remove this and it seemed impossible, but it turned out the replacement battery comes with that part so actually I was supposed to leave it alone.

      step 5: identify the battery

      In the picture of the battery in Step 3, you can see it says something like "3" and "LI???77". There's a piece of metal that I think is welded or something to the top of the battery. It seemed impossible and also maybe not smart to try to remove so I wasn't sure how to find out what an "LI????77" was or how to order another one.

      I've been trying to avoid using LLMs (though I will not get into that because I am exhausted by LLM discourse and I'm sure you are too), but I really had no idea how to figure out what the battery was so I asked an LLM. It gave the response "LIR2477", which (when I looked it up) looked exactly the same as my battery so I figured that was plausible.

      I would be interested to learn non-LLM ways to figure this out though. There must be a way. Lauria showed me how to use DigiKey's search which was very cool though DigiKey didn't have that part.

      (edit: someone in the replies told me that this kind of coin cell battery is named according to its dimensions, and you can use plastic calipers to measure the dimensions of the battery. So I guess a non-LLM way would be to measure the battery with calipers and try to match it to something on the List of battery sizes Wikipedia page, though that page only mentions CR 2477 and not LIR 2477. It's a good example of what's fun for me about trying to avoid LLMs, this "List of battery sizes" page is super interesting and if I use an LLM I might never find it)

      step 6: buy the battery

      I went to AliExpress and ordered:

      1. 2 batteries (I had 2 bike lights and I wanted to fix them both)
      2. some silicone glue to glue things back together

      I think the batteries were $3 each and the glue was $8.

      step 7: solder the new batteries in and glue it back together

      The parts took maybe 2 weeks to arive, and once they arrived, I went back to the makerspace and:

      • soldered in the new batteries
      • put the screws back in. The screws were very small and hard to hold, so at this point I dropped some screws on the ground and couldn't find them because they were too small. So I just used fewer screws and hoped for the best.
      • used the glue to try to put everything back together.
      • Make a somewhat halfhearted attempt to clamp the parts I was gluing together

      Then after waiting some amount of time for the glue to dry I took it home and waited 24 hours for the glue to cure.

      Also I took the old batteries to somewhere nearby that accepts old batteries.

      it works!

      The lights work! I have used them to bike at night! I still haven't needed to recharge them (and tragically I had to order a new Mini USB cable because I got rid of all my Mini USB cables, so I'm still waiting for that), so I still don't know for sure how long the lifetime of the new battery will be.

      Here's what the light looks like after re-gluing. You can see that I didn't glue very carefully. It didn't really go back together that well but I'm hoping it'll be good enough.

      I thought it was really cool that I was able to do this with extremely minimal electronics skills! It cost about $20 CAD to buy the parts, and (whether or not the repair holds up, I'll try to update this post in the future!), it was fun to try to repair something and learn something new.