🏡


  1. October 04, 2026
    1. 🔗 19h/chernobog v6.3.1 release

      Full Changelog : v6.3.0...v6.3.1

    2. 🔗 earendil-works/pi v1.0.2 release

      New Features

      • Sampling by thinking level — samplingParamsByThinkingLevel in models.json sets sampling parameters such as temperature and top_p for each thinking level on OpenAI-compatible APIs. See Configure sampling by thinking level.

      Added

  2. October 03, 2026
    1. 🔗 Simon Willison We're going to need default hard budget caps on pretty much everything rss

      Here's a product feature which the world is going to need a whole lot more of over the coming months and years: default hard budget caps. I'm talking about the feature of pay-by-usage services and APIs that lets you say "after $X/month, cut this thing off and return errors". These need to be hard limits. Soft caps, "after $X/month, send me a warning email", will not cut it.

      Coding agents, and personal agents (coding agents wrapped in a less threatening UI), greatly reduce the friction of spinning up code that can do useful things. Sometimes those things cost money - calls to paid APIs, or hosted web applications, or systems that can bill for additional storage and compute.

      Nobody wants to wake up to an email sent at midnight warning about a budget limit and find that, while they slept, their rogue service had consumed several hundred (or several thousand) more dollars of usage.

      An argument against this is that businesses don't want their hosted applications to start throwing errors because some budget was exceeded. I expect that most businesses and individuals would prefer errors to a surprise $10,000+ bill.

      I think hard budget caps need to be the default. If someone wants to live dangerously they should be able to do that, but it needs to be on an opt-in basis. Have a nice, clear checkbox somewhere prominent:

      Remove the budget cap. My application will not be shut down if I exceed the configured budget limit, and I will be responsible for subsequent charges.

      The service I most want to see this from is AWS. I've heard plenty of stories from people who refuse to use AWS for personal projects out of (justified) fear that a runaway service might bankrupt them. I've also heard stories from people who didn't anticipate this and ended up seriously burned.

      ... and it turns out AWS finally launched spending limits a few weeks ago! From their announcement New AWS experience helps builders get started and ship faster on 16th September:

      When you're ready to upgrade to a paid plan, you can set a monthly spend limit for your project based on your usage patterns so that you stay within your budget. If a project's usage reaches its spend limit, your project is paused for that month.

      See also Create a spend limit in AWS Settings, though that page warns that "We're currently releasing our new experience to a limited number of customers." Here's hoping that hits general availability for existing accounts soon.

      Google Cloud launched a similar feature in July, called Spend Caps, which lets you "set a monthly financial cap on specific services within a project". Looks like this is becoming a trend!

      In an ideal world, our agents could help with this. It would be great if agents started biasing towards recommending providers with hard budget caps, and warning new and inexperienced builders against deploying applications using uncapped services that might get them into trouble.

      You are only seeing the long-form articles from my blog. Subscribe to /atom/everything/ to get all of my posts, or take a look at my other subscription options.

    2. 🔗 anthropics/claude-code v2.1.289 release

      What's changed

      • Fixed a deny or ask rule on a nested part of a compound shell command not holding over a user-installed mod's approval on managed machines
      • Fixed the terminal freezing on short code blocks with many unclosed <script> tags or deeply nested ${ substitutions
      • Fixed Read deny rules not applying to files @-mentioned, changed, or selected in the IDE through a symlink
      • [VSCode] Reverted a 2.1.288 change to claude auth status that may have made sign-outs more frequent
      • Improved how quickly large files open in a plugin code pane by laying the highlighted view out once at its final width
      • Fixed plugin list, plugin eval and plugin update showing a stale copy of a plugin installed from a local folder marketplace, and hot reload for a symlinked --plugin-dir
      • Fixed installed mods not loading in the first session after an upgrade
      • Fixed a plugin's rows above the prompt showing a stale row while the Background tasks dialog was open in fullscreen
      • Fixed plugin panes drawing nothing when a link used a localhost address, an @ in its path, an uppercase host or a file: path
      • Fixed a user-installed plugin being able to rewrite the descriptions of an organization-managed MCP server's sign-in tools
      • Fixed a freeze or forced quit at launch when a plugin drew a Box with a border style the terminal does not know
      • Fixed supervised and background sessions ending when a plugin's on-screen handler threw asynchronously
      • Fixed sessions ending with an interface error when a plugin region with no height kept growing
      • Fixed Bash deny and ask rules missing a command behind an environment variable prefix with an expanded value (e.g. TZ="$HOME" rm -rf build) when the sandbox auto-allows commands
      • Fixed a Bash deny or ask rule being skipped under sandbox auto-allow when a bare variable assignment came before the command
      • Fixed claude plugin validate skipping the plugin when the folder also holds a marketplace manifest
      • Added agent.spawn for teammates, one agent id across plugin hook events, and idle and waiting states in $.agent.list()
      • Fixed sessions ending with "unrecoverable interface error" when a value a mod's ui.render hook wrote made a row throw while drawn; the engine now draws its own row instead
      • Fixed text with a tab, a stray escape and a C1 control, or a short text with a tab and CRLF line endings, drawing over the rows below it
      • Fixed right-aligned content in a mod's pane or band drawing under the close mark or [-], which now also keep one column in from the terminal's edge
      • Fixed a mod's Client that fails while drawn taking down everything the mod drew around it; it now fails alone and raises ui.fault
      • Fixed claude plugin validate failing an Anthropic marketplace's own plugin and listing a clean plugin.json in --json
      • Fixed a mod's band that fails to draw briefly telling the cards under it to step aside
      • Fixed a failed plugin component showing Error or nothing as its reason when the failure carried no message
      • Improved the line a mod's author sees when its band or pane fails to draw: it names the mod and says nothing was drawn
      • Fixed published artifact pages freezing or crashing the reader's browser tab on short code blocks with many unclosed <script> tags
      • Fixed a mod's Client region staying failed for the whole session after the terminal threw while drawing it
    3. 🔗 HexRaysSA/ida-mcp v20261003.0.1 release

      What's Changed

      • Resolve Codex open_database paths from workspace by @thebabush in #30

      Full Changelog : v20260930.0.1...v20261003.0.1

    4. 🔗 HexRaysSA/ida-nexus v0.13.3 release

      What's Changed

      • Improve reference ranking with field-weighted BM25 by @thebabush in #63

      Full Changelog : v0.13.2...v0.13.3

    5. 🔗 earendil-works/pi v1.0.1 release

      New Features

      • Nix flake — nix run github:earendil-works/pi/stable runs the latest release, and nix profile add github:earendil-works/pi/stable installs it. See Install pi.
      • Project overrides for MCP servers — .pi/mcp.json and /mcp can enable, disable, or change the exposure of a user-level server for one project. See Configure servers.
      • MCP Client ID Metadata Documents — oauth.clientRegistration: "cimd" lets authorization servers allow pi by its document URL instead of dynamic registration. See Authenticate with OAuth.
      • Tool renderers for any tool — pi.registerToolRenderer() draws calls to tools that are not registered yet, such as MCP tools in resumed sessions. See Tool rendering.
      • Cloudflare Clef classifiers — @cf/cloudflare/clef and @cf/cloudflare/clef-flash are usable from codemode scripts and extensions. See Use classifier models.

      Added

      • Added a copy key (app.message.copy, default ctrl+x) to OAuth sign-in screens in /login, /mcp, and /mcp login, which copies the sign-in URL when the browser cannot be opened or the wrapped link cannot be selected.
      • Added oauth.clientRegistration: "cimd" for MCP servers, which identifies pi with its Client ID Metadata Document on pi.dev instead of dynamic client registration, so authorization servers can allow pi by URL (#10302)
      • Added project overrides for user-level MCP servers: a .pi/mcp.json entry without command or url sets only enabled, exposure, and toolExposure of the user-level server, and /mcp can enable or disable a server for the current project (#10277)
      • Added Cloudflare's Clef and Clef Flash classifier models to cloudflare-workers-ai, usable from codemode scripts and extensions (#10316 by @ndisidore, #10322 by @RealAlexandreAI)
      • Added pi.registerToolRenderer(), which chooses how calls to a tool are drawn, including tools that are not registered (#10285)
      • Added a Nix flake for macOS and Linux: nix run github:earendil-works/pi/stable runs the latest release, and nix profile add github:earendil-works/pi/stable installs it. See Install pi (#9137)

      Changed

      • pi update on global npm installations now recommends migrating to the managed installation from the pi.dev installer, which pins all dependencies.
      • Anthropic tools added or redefined mid-conversation are now defined inline in the conversation, so redefining a tool under the same name keeps the prompt cache instead of resending the full tool list.

      Fixed

      • Fixed installations resolving vulnerable brace-expansion 5.0.9 by pinning brace-expansion 5.0.12 as a direct dependency (GHSA-q2hr-2g5m-vwhr, GHSA-qhr7-859c-m2p7, GHSA-6j4f-fj2g-mc7p) (#10288)
      • Fixed a trailing comma in --models adding an extra model to the model cycle (#10334)
      • Fixed a codemode script that prints in a loop crashing pi by running out of memory: a script fails once its output passes 16 Mi characters or 100000 items (#10283)
      • Fixed JPEG, GIF, and WebP images rendered by extensions through Image not appearing in Kitty, Ghostty, WezTerm, and Warp (#10292)
      • Fixed MCP tool calls in resumed sessions and HTML exports rendering fully expanded until their server connected, or for good if it never did (#10285)
      • Fixed fullscreen Kitty images collapsing to a one-row strip after scrolling in WezTerm (#10319)
      • Fixed "Selected model is at capacity" provider errors ending the turn instead of being retried (#10278)
      • Fixed Cloudflare AI Gateway Claude models failing with a 404 by using dashed model IDs (claude-opus-5-5 instead of claude-opus-5.5)
      • Fixed Sign in with ChatGPT continuing when its callback port is taken by another login, which made the browser show "OAuth state mismatch"; it now fails with a port-in-use error (#10265)
      • Fixed Amazon Bedrock OpenAI models costing requests above 272k input tokens at the short-context rate; Bedrock models now include the pricing tiers listed on models.dev (#10326)
      • Fixed Amazon Bedrock Claude requests failing with "Invalid signature in thinking block" after the system prompt or tools changed (#10324)
      • Fixed Together DeepSeek V4 Pro losing its thinking level controls after Together renamed it to deepseek-ai/DeepSeek-V4-Pro-0813 (#10336 by @cv)
      • Fixed the default NVIDIA model pointing at nvidia/nemotron-3-super-120b-a12b, which NVIDIA no longer serves; the default is now nvidia/nemotron-3-ultra-550b-a55b

      Removed

      • Removed npm-shrinkwrap.json from the published package. npm installations no longer pin transitive dependencies, and library consumers can now override them. Use the pi.dev installer for pinned installations (#5653)
    6. 🔗 r/LocalLLaMA Yes bots we get it, Strata is good now please stop rss

      Yes bots we get it, Strata is good now please stop | It's like the entire sub has become that scene from Konosuba where the cult keeps making up fake scenarios saying the only solution is to join their religion submitted by /u/Mayion
      [link] [comments]
      ---|---

    7. 🔗 r/LocalLLaMA Aleph-Alpha/Kolibri-1 · Hugging Face - 78B parameters. 3.46B active. Up to 1M tokens of context - Apache 2.0 rss
    8. 🔗 r/LocalLLaMA I'm pretty close to the middle thanks to you all rss

      I'm pretty close to the middle thanks to you all | It's been a blast and learning a ton. But seriously, you all have me down a rabbit hole that my wallet and hours of sleep need to be pulled out of. submitted by /u/pmarsh
      [link] [comments]
      ---|---

  3. October 02, 2026
    1. 🔗 MetaBrainz Our editors: vzell rss

      This is part of a series where we talk to some of the amazing editors that have cracked one million MusicBrainz edits and qualified for super special (lousy) shirt plan.*

      These users have donated substantial amounts of their time to improve open music data, enriching open-source knowledge and culture for billions of people. Let’s find out what makes vzell click. And click and click and click!

      Thanks so muchvzell for agreeing to chat. Many MusicBrainz editors will know you through your MusicBrainz profile and editing - who are you offline?

      Hi, I'm from Germany (originally born in Romania, Transylvanian of German descent). I live in Drabenderhöhe—the largest Transylvanian-dominated community near Cologne—where we keep all Transylvanian traditions alive.

      I am an avid singer myself and a member of an eight-time championship-winning choir; it is my favorite hobby, right after Bruce Springsteen and MusicBrainz. I studied physics and have been retired for a year, yet I still work three days a week in the IT industry.

      How did you get started with editing MusicBrainz?

      It's all because of Bruce Springsteen dominating my life musically. As a data fetishist, I want to record every conceivable piece of information in a super database like MB.

      Ah, a MusicBrainz obsession sprouting from another music obsession. That will be familiar to many MusicBrainz editors.

      How does editing MusicBrainz fit into your everyday life and routine?

      Perfectly… as I sit in front of the computer anyway, listening and archiving Bruce and related stuff.

      Do you have any personal tricks or rules to keep a large amount of editing sustainable, healthy and fun?

      Yes… have a lot of friends to party with, a choir to sing with AND listen to Bruce during editing (a never ending story, this guy has so many bootlegs and I have them all, 30 TB on my NAS). Last but not least, stay focused and work precisely, it's worth it… and… laugh and enjoy life, there seems to be actually only one.

      Do your friends and/or family know that you edit MusicBrainz for (hopefully) fun? What do they think about it?

      All of them know about it and now with the T-shirt it will wide spread even more 🙂

      Actually they like it, and they know it's my habit to work precisely no matter what it is that I'm doing.

      What is your favourite thing to edit?

      Right now, entering every single Bruce bootleg on earth after I finished off (a long time ago) entering all his events (about 4,200) and works and having a lot of beer…

      What is your least favourite thing to edit?

      Bruce related Various Artist albums, but I do it anyway…

      What are your favourite editing tools? Do you have any editing tips?

      In the beginning my fingers, later on I automated as much as possible with self written Emacs Lisp routines (yes I'm an Emacs guru). Nowadays userscripts I've written with the help of AI (Claude), see [forum thread].

      One tip so, be consistent in what you're doing, every single dash/en-dash/em- dash matters… hahahaha

      Last but not least, what are your current top music recommendations!

      Well, do yourself a favour and listen to Willy DeVille and maybe… hahahaha… to our choir MGV Drabenderhöhe.

      And last but not least "The Boss"…

      Thanks again vzell. Happy editing!

      *If you are a million-editor and missed our emails and are keen for a shirt and/or a chat, please get in touch. If you have a million edits you should know how to reach staff, so no contact link for you! Non-million editors, you are also welcome to share your story, please get in touch (you get a link!)

    2. 🔗 r/LocalLLaMA Buying RTX 5090 At Micro Center Reportedly Now Requires Paperwork, Including A No-Export Declaration rss

      Buying RTX 5090 At Micro Center Reportedly Now Requires Paperwork, Including A No-Export Declaration | submitted by /u/Boomfrag
      [link] [comments]
      ---|---

    3. 🔗 anthropics/claude-code v2.1.288 release

      What's changed

      • Added $.ui.selection() for mods: returns the text you last selected in fullscreen mode and, when the selection lies within one transcript row, that row
      • Added a built-in gh api to cloud sessions whose image has no GitHub CLI, and fixed the built-in sending control characters from file names, jq filters or GitHub errors to the terminal
      • Added recovery for a prompt cleared with Ctrl+C: pressing Up on the empty prompt brings the draft back, including pasted text and images
      • Added a re-authenticate prompt when an MCP server asks for more OAuth scope during a tool call
      • Added --max-findings <n>|all to /code-review to report more or fewer findings than the usual limit; the choice is reused until you pass --max-findings default
      • Added Ctrl+F to find a session by name and Alt+↑/↓ to jump between groups in the agents view; both, and rename, can be rebound in keybindings.json
      • Added a screen reader mode announcement of the new permission mode when you approve a plan, including with Shift+Tab
      • Fixed mid-response API timeouts failing the turn: non-interactive sessions and subagents now continue from the partial response, and thinking-only responses are retried
      • Fixed long conversations failing with "Prompt is too long" instead of auto-compacting when the last reply reported zero token usage
      • Fixed --resume sometimes dropping files and other context that a compaction had just restored
      • Fixed a resumed session sometimes not saving the last response of a turn, so that the next --resume showed the prompt unanswered
      • Fixed resume occasionally loading a transcript cut short when the same session rewrote the file during the load
      • Fixed resuming a conversation started on 2.1.286 or earlier dropping the model's earlier thinking
      • Fixed session titles, memory recall and prompt hooks failing on Mantle or behind gateways that reject structured outputs; added CLAUDE_CODE_DISABLE_STRUCTURED_OUTPUTS to turn structured outputs off
      • Fixed auto mode denials pointing Claude at a Bash permission rule when the blocked tool was not Bash
      • Fixed auto mode on Bedrock and Mantle switching to the local classifier for the rest of the session after a request to an older model, such as a WebFetch summary or a sonnet subagent
      • Fixed cloud sessions that restarted on a newly picked model replying with that model after the server refused it
      • Fixed Cowork cloud sessions staying marked as waiting for input after a WebFetch permission prompt for an unapproved URL went unanswered for five minutes
      • Fixed prompt suggestions not appearing on a phone that joins a Cowork cloud session started on another device
      • Fixed a mod's button sometimes running a different button's action when pressed on a view drawn before Claude Code restarted
      • Fixed a plugin's pane showing nothing when one Code element held a diff that does not parse; it now draws as plain code
      • Fixed plugin LSP servers receiving literal ${user_config.*} and ${CLAUDE_PLUGIN_ROOT} placeholders in initializationOptions and settings instead of substituted values or manifest defaults
      • Fixed a plugin's tool.call hook making Bash fail and file searches read the wrong folder in subagents that run in a worktree
      • Fixed git-subdir plugin installs failing, or caching an incomplete plugin, on older git (before 2.39, e.g. Ubuntu 22.04's 2.34)
      • Fixed plugins loaded with --plugin-dir not showing "Configure options" in /plugin
      • Fixed background sessions ending when a plugin was reloaded or disabled while one of its timers or reads was still running
      • Fixed sandboxed heredocs with an unquoted delimiter (`python3 <
    4. 🔗 Kagi An Update on Orion for Linux and Windows rss

      Kagi is ending development of Orion for Linux and Windows and open-sourcing both so the community can carry them forward. Our small team will now focus fully on making Orion for macOS and iOS faster, more stable, and more capable.

    5. 🔗 r/LocalLLaMA I made my iPhone a second GPU for my 24 GB MacBook: Qwen 3.8 27B prefills 29–44% faster & my holds part of the CTX window. rss

      I made my iPhone a second GPU for my 24 GB MacBook: Qwen 3.8 27B prefills 29–44% faster & my holds part of the CTX window. | DISCLAIMER THE PREFILLING TPS SHOWN ON THE PHONE IS COMPUTED ONLY FOR THE LAYERS IT HOLDS. ALREADY FIXING IT TO SHOW END-TO-END PREFILL RATE. NUMBERS BELOW ARE ACCURATE FOR E2E PREFILL RATE. Every file or tool result my agent reads on a 24 GB M4 Pro MacBook is a wait, and 64k of 8-bit context is all that fits next to Qwen 3.8 27B (IQ4_XS), even with the wired limit raised to 20480. An iPhone 17 Pro Max was sitting in my pocket, so I figured what can I do to make use of this extra silicon. Turns out a 10 Gb/s USB-C cable & some software is all you need. The Mac runs layers 1–40 of each 256-token batch and streams the activations to the phone. The phone runs layers 41–64 on its GPU while the Mac starts the next batch. The A19 Pro's GPU has matrix units (Metal 4 tensor ops), and they make the phone's half 2.4x faster than the same phone without them. Same build, phone off vs. on, prefilling a 2,000-token file into a saved agent session:

      • 8k context: Mac alone 132 tok/s → Mac + iPhone 177 tok/s (+35%) (measured two days earlier, same bench)
      • 16k context: Mac alone 109 tok/s → Mac + iPhone 157 tok/s (+44%)
      • 32k context: Mac alone 101 tok/s → Mac + iPhone 130 tok/s (+29%)
      • 48k context: Mac alone 87 tok/s → Mac + iPhone 113 tok/s (+30%)

      A fresh 27k-token agent session, cold: 245 s on stock llama.cpp, 228 s on my fork with the Mac alone, and 168 s with the phone. Past 64k the phone switches jobs. The oldest KV pages move to the phone and the Mac runs all 64 layers. For every attention layer, the phone computes attention over the old keys on its GPU, and the Mac merges that with its own part. While writing, the phone's Neural Engine takes part of that work too: each 16k-key page of old context is compiled into a Neural Engine model with the keys as its weights. At 140k that took writing from 279 to 176 ms per token compared with the phone's GPU alone. The server allocates 196k–229k of 8-bit context based on the phone's free memory; that's up to ~5.7 GB of KV cache living on the phone instead of the Mac, so the Mac's memory use stops growing at 64k. I've tested a growing session to 128k at 8-bit, with 3/3 planted facts recalled. Separately, at 140k in 4-bit, the run passed the gate with greedy output matching the Mac-only run for 32 generated tokens. What it doesn't do: speed up writing below 64k. That's the Mac's job. My fork's kernels (SME2 on the M4 CPU and Metal fusions) plus DFlash2 speculative decoding take it from 11.3 tok/s on stock llama.cpp to 25 tok/s at about 30k context with medium thinking, phone or not. SME2 also adds up to 29% to prefill on the Mac alone. Past 64k the phone does share the writing (attention over the old keys), and without it the Mac would have to drop to 4-bit context to reach 128k. In real use I have seen upwards of 30 TPS at lower context. The phone joins prefills over about 512 tokens. In one real session, that was 7 of 36 requests, but about 83% of the tokens read. Past 64k it holds the context and does the old-key attention, but it stops running layers 41–64 there for now; doing both is next. One request at a time. I'm curious what this setup could do with newer model architectures. DeepSeek V4.1-Flash reports 890 bytes per token for its global KV cache and adds n-gram embedding tables (Engram). Qwen3.8-Flash-Next, the Qwen 4 architecture preview, has an n-gram lookup table too. Those aren't features of the 27B model I tested, and I haven't benchmarked either architecture here. The real gold is within the newer phones and models working together. With the A20 Pro in the iPhone 18 Pro Max, I bet there is a lot more for me to push. Code, setup and bench scripts: https://github.com/StayLameBro/backburner Still a lot of work to do but I built this with Opus 5.5. Happy to answer anything. submitted by /u/StayLameBro
      [link] [comments]
      ---|---

    6. 🔗 r/LocalLLaMA Qwen3.8-27B-Humanlike-Chat 2.0: texts like a human, now with tool calls and better instruction following rss

      Qwen3.8-27B-Humanlike-Chat 2.0: texts like a human, now with tool calls and better instruction following | Last month I posted a Qwen3.8-27B LoRA that makes it talk like a person instead of an assistant. It got a lot more attention than I expected: 700+ upvotes, 248 comments and 44k downloads since. I read every comment. People really don't like assistant speak, so its tone of voice resonated. The rest got roasted, very fairly:

      incapable of producing more than a few words at a time. single default personality which no amount of prompting can overcome will not use tools , at all, whatsoever. There needs to be a middle ground

      They were right. The tool calls didn't actually work, and when people asked it to do something it would sometimes just say it's busy or going to bed. Very human. In a bad way. So I spent the last three weeks on 2.0. The goal was simple: keep the voice people liked and lose the drawbacks. What 2.0 does now

      • With no system prompt, it's a normal person texting. Not an assistant, not a catgirl.
      • Give it a character card and it becomes that person, and still texts like one.
      • Ask for a formal email, numbered steps or a proper explanation, and you get exactly that. Then it goes back to texting.
      • Don't want the lowercase texting? Tell it "from now on write in full sentences" (or put it in the system prompt) and it sticks to that until you say otherwise. v1 ignored this completely.
      • It calls tools, and it asks when something is missing instead of making it up. This is the part I'm happiest about. Ask the base model to book a flight without saying where from and it picks JFK. 2.0 asks where you're flying from.
      • It writes code and does math at roughly base-model level.

      It's a colleague and a humanlike companion, not an assistant. Use it for chat, roleplay, agents or actual work. How I trained it v1 was plain SFT on real and synthetic conversations (139,845 messages from 1,396 conversations). That copies habits, including the bad ones. For 2.0 I used on-policy distillation. The model writes its own replies and a teacher grades every token. There are two teachers:

      • v1 plus a hidden "text like a person" instruction, for chat and characters;
      • the plain base model, for instructions, tools and code.

      The student never sees the hidden instruction, so it learns the behaviour without needing a prompt. Same 27B, a second LoRA on top, merged. Numbers (vs the model I trained on, huihui-ai's abliterated Qwen3.8-27B; same prompts, same run, thinking off) | Benchmark | Base (abliterated) | 2.0
      ---|---|---
      IFBench (instruction types I never trained on) | 37.3 | 43.7
      When2Call (call, ask or refuse correctly) | 48 | 58
      BFCL irrelevance (don't call a tool when none fits) | 60 | 78
      IFEval, GSM8K, BFCL simple | 81.9 / 89.1 / 97 | 83.5 / 89.1 / 98 (ties)

      Full chart in the images.

      Where it's still worse: knowledge (MMLU-Pro 72.5 vs 78.5) and competitive code (LiveCodeBench 51 vs 56).

      Is it actually more human? I built a benchmark for this, "ishuman":

      • It takes 150 fragments from unseen chats.
      • Has each model write the next message.
      • Shows a judge the real message and the model's without labels, and asks which one a person wrote.

      Model | Judge thought it was the real person (50% = can't tell)
      ---|---
      Qwen3.8-27B abliterated (huihui-ai, the model I trained on) | 0.3%
      Same abliterated model + a "text like a human" system prompt | 6.8%
      Qwen3.8-27B official (unmodified, via OpenRouter) | 15.1%
      Qwen3.8-27B-Humanlike-Chat 2.0 | 23.5%

      So no, you can't just prompt your way there. In a separate test of 16 live multi-turn chats with invented people, 2.0 was picked over the base model 16 out of 16 times.

      Links

      Big thanks to everyone who left feedback last time, especially the ones who were critical. Tell me where it still sounds like an assistant.

      Edit: safetensors are up for vLLM and SGLang:
      GPTQ-Int4 (24 GB): https://huggingface.co/LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0-GPTQ- Int4
      FP8 (48 GB): https://huggingface.co/LessThanThreeAI/Qwen3.8-27B-Humanlike- Chat-2.0-FP8
      BF16 (80 GB): https://huggingface.co/LessThanThreeAI/Qwen3.8-27B-Humanlike- Chat-2.0

      submitted by /u/kvyb
      [link] [comments]

    7. 🔗 r/LocalLLaMA New in llama.cpp: Decision Models rss

      New in llama.cpp: Decision Models | submitted by /u/paf1138
      [link] [comments]
      ---|---

    8. 🔗 Luke Muehlhauser Media diet for Q3 2026 rss

      Music

      Music I most enjoyed discovering this quarter:

      This quarter, I switched my focus from classical (including contemporary classical) music to jazz, because I hit diminishing returns on finding classical composers I liked. New-to-me jazz artists I "strongly liked" some music from were:

      • Phronesis: most of Organic Warfare (2007), some of Green Delay (2009), some of Walking Dark (2012), most of Life to Everything (2014), some of Parallax (2016), some of The Behemoth (2017), "One for Us" & "The Tree Did Not Die" (2018)
      • Jasper Hoiby: "Fellow Creatures" (2016), some of Conversations of Hope (2026)
      • Gard Nilssen: some of If You Listen Carefully the Music is Yours (2020), "The Space Dance Experiment" & "SP68" (2023), some of Great Intensions (2025)
      • Guillaume Perret: most of Guillaume Perret & the Electric Epic (2012), "Doors" & "Shoe Box" (2013), most of Open Me (2014), "She's Got Rhythm" (2016), some of A Certain Trip (2020)
      • Andy Emler: Mega Octet "Part 1" & "Part 7" (1990), Head Games "Part 6" (1992), "Bacteria" (2003), "Urbanhof" & "Minicrobe" (2004), some of West in Peace (2007), most of Crouch, Touch, Engage (2009), some of E Total (2012), "Finally Closing" (2014), most of Obsession 3 (2015), most of A Moment For (2018), some of No Rush! (2021), some of The Useful Report (2022), some of Le Temps est parti pour rester (2025)
        • With these new listens, Emler has crossed the 5-hour mark as one of my favorite musical artists!
      • Ola Kvernberg: "Bricks of Stone" (2006), "Roland" (2009), some of Liarbird (2011), most of The Mechanical Fair (2014), some of Northern Tapes (2014), some of Steamdome (2017), most of Steamdome III (2024), "Let's Go" & "Ludvig Takes Over" (2026)
      • Ivo Neame: some of Strata (2015), "The Rise of the Lizard People" & "Strega" (2021)
      • Fred Pallem: "Ti seppellirei neil giardino" & "Mustang Pursuit" (2011), some of Le retour! (2015), "Get It In Orbit" & "Bitches en Marbella" (2022), "Poursuivi par des elephants genats" (2022)
      • Lukas Kranzelbinder / Shake Stew: "Milking of the Mugwumps" & "Dreams of Sex and Crime" (2012), "Rise of the Black Centipede" & "Hans" (2015), most of The Golden Fang (2017), most of Rise and Rise Again (2018), most of Gris Gris (2019), "Play Mass" (2020), most of Heat (2022), "Lila" & "Shasta Fey" (2023), most of Ten One Two (2026)
        • With these new listens, Kranzelbinder has crossed the 5-hour mark as one of my favorite musical artists!
      • Zbigniew Seifert: "Trubulent Plover" & "On the Farm" (rec. 1976), "Impressions" & "Spring on the Farm" (1979), "Passion" & "Pinocchio" (1979)
      • Chris Lightcap: some of Deluxe (2010), most of Epicenter (2015)
      • LBT: some of Levitation (2016), Way Up in the Blue (2018), most of Stereo (2020), Make Kin (2021), most of Abstrakt (2023), House (2023), "Jabal" (2023)
      • Frederic Maurin / Ping Machine: some of Encore (2013)
      • Magic Malik: some of 69-96 (2000), "Passage a vide" (2002), "XP 21" (2004)

      I also listened to a significant portion of the recorded works by each of the (new-to-me) "classical" composers listed below.1 My favorites pieces from them (names linked to playlists) were:

      • Johan Halvorsen (b. 1864): "Entry March of the Boyars" (1893)
      • Geronimo Gimenez (b. 1854): El baile de Luis Alonso "Preludio" & "Intermedio" (1896), La boda de Luis Alonso "Intermedio" (1897)
      • Leo Weiner (b. 1885): Hungarian Folk Dances Suite mvts 1, 2, 4 (1931)
      • Plus the following composers for which I didn't "strongly like" any of their pieces I listened to: Jan Novak (b. 1921), Fisher Tull (b. 1943), Ketil Hvoslef (b. 1939), Ma Sicong (b. 1912), Jef Penders (b. 1928), Robert Muczynski (b. 1929), Francisco Esteve Pastor (b. 1915), Manuel Carrascosa García (b. 1911), Antonio Soler (b. 1721), Karl King (b. 1891), Henry Fillmore (b. 1861), Manuel Lopez Farfan (b. 1872), Boris Tchaikovsky (b. 1925), Peter Kleine Schaars (b. 1962)

      Rediscovered or revisited, and strongly liked:

      • Hidden Orchestra: some of Night Walks (2010), some of Archipelago (2012), some of Dawn Chorus (2017)
      • Brandt Brauer Frick: "You Make Me Real" & "Mi Corazon" (2011), "Ocean Drive (Schamane)" & "Skiffle It Up" (2013)
      • Ozric Tentacles: some of There Is Nothing (1986), some of Sliding Gliding Worlds (1988), Pungent Effulgent (1989), Erpland (1990), Strangeitude (1991), Jurassic Shift (1993)
      • Spring Heel Jack: 68 Million Shades (1996), most of Busy, Curious, Thirsty (1997), most of Treader (1999), Bombscare (2000), Disappeared (2000), "Chorale" & "Chiarascuro" (2001), Amassed (2002), "Lata" & "Autumn" (2004)

      And for jazz specifically:

      • Esbjorn Svensson / E.S.T.: "Dodge the Dodo" (1999), some of Seven Days of Falling (2003), "The Unstable Table & the Infamous Fable" & "A Picture of Doris Travelling with Boris" (2005), some of Tuesday Wonderland (2006), "Three Falling Free, Pt. II" (2012)
      • Hedvig Mollestad: "Antilone" & "Ekhidna" (2020), some of Maternity Beat (2022)
      • GoGo Penguin: most of v2.0 (2014), some of Man Made Object (2016), some of A Humdrum Star (2018), some of Ocean in a Drop (2019), most of GoGo Penguin (2020)
      • Martin Kuchen / Angles: "By Way of Deception" (2012), "European Boogie" & "Ubabba" (2014), "Equality & Death" & "Ardor" (2017), "The Hidden Balcony" & "Fkk Down, Fkk Off" (2022)
      • Snarky Puppy: "Sintra" & "Gretel" (2015), “Trinity” (2022), "Lingus" (2014), "It Stays With You" (2025)
      • Wacław Zimpel: "Afterimages" (2013), some of LAM (2016), Zimpel/Ziołek (2017)
      • Brad Mehldau: "Luxe" (2014), "The Prophet is a Fool" (2019), "Herr und Knecht" (2022)
      • Jaga Jazzist: most of A Living Room Hush (2001), some of The Stix (2003), most of One-Armed Bandit (2010)
      • VSOP the Quintet: most of V.S.O.P. The Quintet (1977), most of Tempest in the Colosseum (1977), most of Five Stars (1979), most of Live Under the Sky (1979)
      • Isfar Sarabski: most of Planet (2021)
      • Jaimie Branch: some of Fly or Die (2017), "Simple Siver Surfer" & "Nuevo Roquero Estereo" (2019), most of Fly or Die Fly or Die Fly or Die (World War) (2023)
      • Abercrombie / Holland / DeJohnette: some Gateway (1975), "How's Never" (1995)
      • Larry Coryell: Barefoot Boy (1971)
      • Marvin "Hannibal" Peterson: "Song of Life" (1974), some of Hannibal (1975), "Now Stand" (1978), some of The Angels of Atlanta (1981)
      • Jane Ira Bloom: "Overstars" (1987), some of Art & Aviation (1992)
      • Bobby McFerrin: "He Ran All the Way" (1990), Circlesongs (1997), some of VOCAbuLarieS (2010)
      • John McLaughlin & Shakti: "Joy" (1976)
      • David Friesen: some of Star Dance (1976), "Waterfall Rainbow" & "Dancing Spirits Before the Lord" (1977)
      • Keith Tippett: some of You Are Here… I Am There (1970), some of Dedicated to You, but You Weren 't Listening (1971)
      • Ernie Krivda: some of Satanic (1977), some of The Alchemist (1978)
      • Wynton Marsalis: “Bullet Train” (1999), Swing Symphony mvts 3, 4 (2019)
      • Møster!: "Underworld Risk" (2014), some of When You Cut into the Present (2015), "Unhorsed by Chivalry" (2018), some of Dust Breathing (2020), "The Electric Wood Orchestra / Dreaming Xaxado…" (2024)
      • Ponga: Ponga (1999), "Hagro" (2000)
      • Ned Rothenberg: some of Overlays (1991), "Hidalgo" (1995), some of Real and Imagined Time (1995)
      • Marty Fogel: most of Many Bobbing Heads, at Last (1989)
      • David Torn: some of Cloud About Mercury (1987), some of Polytown (1994), some of Tripping Over God (1995), some of What Means Solid, Traveller? (1996)
      • Matthew Shipp: some of Nu Bop (2002), "Cohesion" (2003)
      • Ben Neill: most of Torchtower (1994), Triptycal (1996), "Tunnel Vision (Spring Heel Jack Holland Tunnel Mix)" (1998)
      • James Brandon Lewis: some of The Messthetics and James Brandon Lewis (2024), some of Deface the Currency (2026)
      • The Bad Plus: most of These Are the Vistas (2003), most of Give (2004), some of Suspicious Activity? (2005), most of Prog (2007), some of Never Stop (2010), most of Made Possible (2012), some of Inevitable Western (2014), most of The Bad Plus Joshua Redman (2015), some of Never Stop II (2018), most of Activate Infinity (2019), "Sun Wall" (2022), some of Complex Emotions (2024)
        • With these new and old listens, The Bad Plus has crossed the 5-hour mark as one of my favorite musical artists!
      • Makaya McCraven: "Gnawa" & "On the Spot" (2015), "Atlantic Black" & "Inner Flight" (2018), "This Place That Place" & "Seventh String" (2022)
      • Vienna Art Orchestra / Mathias Ruegg: "Jelly Roll, but Mingus Rolls Better" (1990), some of Art & Fun (2002), "We take Pride…" (2004), "French Alphorn" (2010)
      • Miho Hazama: some of Beyond Orbits (2023)
      • Toshiko Akiyoshi: “Henpecked Old Man [22m version]” (1976), “Minamata” (1978)
      • Paul Winter: "Icarus" (1970), “Whole Earth Chant” (1972), "Kyrie" (1987)
      • Dollar Brand (Abdullah Ibrahim): “Hajj (The Journey)” (1978), some African Marketplace (1980)
      • Eric Dolphy: "Hat and Beard" (1964)
      • Dave Douglas: “Three Beasts” (1997), “Ruckus” and “Witness” (2001), Freak In (2003), Keystone (2005), most of Moonshine (2007), most of Soundtrack (2010), Expand (2010), Burst (2010), High Risk (2015), Dark Territory (2016), "Climate Strike" (2020), most of Transcend (2026)
      • Jean-Luc Ponty:2 "Contact" (1969), "How Would You Like to Have a Head Like That" (1970), Astrorama (1970), some of Sonata Erotica (1976), Upon the Wings of Music (1975), Aurora (1976), Imaginary Voyage (1976), Enigmatic Ocean (1977), most of Cosmic Messenger (1978)

      I also listened to a significant portion of the recorded works by each of the (not new-to-me) "classical" composers listed below.3 My favorites pieces from them (names linked to playlists) were:

      • Raimo Kangro (b. 1949): Concerto No. 2 for Two Pianos and Chamber Orchestra mvts 2, 4 (1988), "Display I: Portrait of Steve Reich" (1991), "Display II: Portrait of Mozart" (1991), Display VIII: Portrait of Schubert (1995), Arcus (1998), Piano Concerto No. 2 (1999)
      • Johann Strauss I (b. 1804): "Sperl Galopp" (1831), "Cachucha Galopp" (1837), "Versailler Galopp" (1839), "Radetzky March" (1848)
      • Scott Joplin (b. 1868): "Maple Leaf Rag" (1899), "The Entertainer" (1902)
      • François Couperin (b. 1668): "The Mysterious Barricades" (1717)
      • Luigi Boccherini (b. 1743): String Quintet G. 275 , mvt 3 (1771), Symphony No. 4 G. 506 , mvt 3 (1771)
      • Vittorio Monti (b. 1868): "Csardas" (1904), "Coquetterie" (1913)
      • Plus the following composers for which I didn't "strongly like" any of their pieces I listened to: Domenico Scarlatti (b. 1685), Hans Werner Henze (b. 1926)

      Movies/TV

      Ones I "really liked" (no star), or "loved" (star):

      • Barker: Obsession (2026) ★
      • Hikari: Rental Family (2025)
      • Mollner: Strange Darling (2023)
      • Trier: Sentimental Value (2025)
      • Johnson: Blackberry (2023)
      • Various: Smiling Friends , seasons 1-2 (2022-2024)
      • Various: Common Side Effects , season 1 (2025)
      • Jones: I Swear (2025)
      • Parsons: Backrooms (2026) ★
      • Moore: The Outfit (2022)
      • Borgli: Sick of Myself (2022)
      • Borgli: The Drama (2026)
      • Stanton: Toy Story 5 (2026) ★
      • Blitz: Review , season 2 (2015) ★
      • Blitz: Review , season 3 (2017)
      • Johnson: Nirvanna the Band the Show the Movie (2025) ★
      • Eggers: Nosferatu (2024)
      • Various: The Larry Sanders Show , season 1 (1992)
      • Ricky Gervais & Stephen Merchant: Extras , season 2 (2007)
      • Brough: Rosehaven , season 1 (2016)
      • Wilde: The Invite (2026)
      • Various: The Comeback , season 1 (2005)
      • Various: The Comeback , season 2 (2014)

      Games

      All games I finished or decided to stop playing:

      • [none]

      Standup comedy

      • Kelsey Cook: Happy Hour (2026)

      Books

      I post book ratings and reviews to my Goodreads account instead of here.

      1. The pieces I listened to for each composer were: Novak, Tull, Hvoslef, Ma, Penders, Muczynski, Esteve Pastor, Carrascosa García, Halvorsen, Gimenez, Soler, King, Gillmore, Lopez Farfan, B. Tchaikovsky, Weiner, Kleine Schaars.
      2. Because so many of my favorite Ponty tracks are covers, in his case I'm strictly limiting my Favorites playlist to non-covers, which I don't do so strictly for other artists' Favorites playlists.
      3. The pieces I listened to for each composer were: Kangro, Strauss I, Joplin, Couperin, Boccherini, Monti, Scarlatti, Henze.
    9. 🔗 r/LocalLLaMA Does anyone know if any new releases from Mistral are planned? rss

      Does anyone know if any new releases from Mistral are planned? | It’s been a long time since the last models came out. I notice they are selling GLM on the site, and I wonder if they are developing something, given the long silence. submitted by /u/LegacyRemaster
      [link] [comments]
      ---|---

    10. 🔗 Confessions of a Code Addict How and Why fork() Uses Copy-on-Write rss

      In the last video, we learned about copy-on-write (CoW) and understood its mechanics by digging inside the mmap pathway in the kernel. This is a continuation of that where we look at the fork() system call and how it uses copy-on-write behind the scenes.

      Following is a timestamped list of chapters in the video to help you navigate quickly. Also, I recommend watching the video at a higher speed to get a better experience.

      • (00:00) Introduction : Copy-on-write with fork() and the Instagram example.

      • (02:54) Copy-on-write recap : Sharing physical frames and making private copies when a process writes.

      • (10:08) What fork() does : Creating a child process and inheriting the parent's state.

      • (13:15) A code example : Return values, inherited memory, and independent writes in parent and child.

      • (18:26) Why use copy-on-write? : The cost of copying the parent's memory.

      • (21:05) Page tables after fork() : How parent and child initially share physical frames.

      • (23:45) The shell and exec() : Why copying memory upfront would often be wasted work.

      • (26:04) Handling a write : Read-only page-table entries, page faults, VMA permissions, and copying a frame.

      • (30:01) Benefits and tradeoffs : Memory savings and the performance cost of deferred copying.

      • (33:41) Instagram 's memory-sharing problem: Shared memory gradually becoming private in Python worker processes.

      • (41:14) Python objects and reference counting : Why an expression like if x is None can cause memory writes.

      • (45:16) Inside the interpreter : How evaluating an expression changes reference counts.

      • (48:12) Immortal objects : Avoiding reference-count updates for objects that live throughout the runtime.

      • (50:19) Wrap-up : Plans for a hands-on demonstration using Linux memory debugging tools.

      If you are new to this series, it is based on my ebook called "Virtual Memory from First Principles". It is available to read for free online and also available to purchase from Gumroad (PDF/Epub) and Amazon (Kindle edition).

      Buy PDF/Epub

      Get Kindle Edition


      And, if you want to watch the previous videos in this series, the following is what has been published so far:

      Share

      Read more

    11. 🔗 smol-machines/smolvm smolvm v1.22.2 release

      What's Changed

      • Follow the fetch_update rename so clippy on current stable passes by @BABTUNA in #1505
      • Let the API amend a stopped machine's egress allow list by @BABTUNA in #1501
      • Allow clippy::double_must_use on the shim's async trait by @LoganGrasby in #1508
      • Reject unknown fields on the mutating API request types by @BABTUNA in #1509
      • Bump the workspace to 1.22.2 by @BinSquare in #1513

      Full Changelog : v1.22.1...v1.22.2

    12. 🔗 HexRaysSA/plugin-repository commits sync repo: -1 plugin, +1 release, -1 release rss
      sync repo: -1 plugin, +1 release, -1 release
      
      ## New releases
      - [ida-nexus](https://github.com/hexrayssa/ida-nexus): 0.13.2
      
      ## Removed plugins
      - ida-codemode
      
    13. 🔗 backnotprop/plannotator v0.27.25 release

      Follow @plannotator on X for updates

      Missed recent releases? Release | Highlights
      ---|---
      v0.27.24 | Image previews stay in the all-files view, PR comment previews open on the commented line
      v0.27.23 | Bitbucket Cloud PR review, Question UI for answering agents in place, opt-in auto-update, viewed files remembered, review another repo or worktree
      v0.27.22 | Plans open the docs they link to, Claude review jobs locked down, Code Tour on Linux without Claude's sandbox, Pi reviews the latest plan
      v0.27.21 | Remote and phone sessions load several times faster, real Request changes on GitHub, model pickers show real names, OpenCode fixes
      v0.27.20 | Mistral Vibe support, annotate gets the full Options menu and Settings, jj Commits panel, long lines wrap in plan code blocks
      v0.27.19 | Before/After image previews in code review, file comments as GitHub file threads, forge-correct #123 links, /plannotator-last finds the right session
      v0.27.18 | Model pickers from your installed Claude and Codex (Opus 5.5, Fable 5.1, GPT-6), unsent PR review comments survive new pushes
      v0.27.17 | Diagram files open in the diagram viewer, OpenCode switches model with agent, idle review stops polling the git remote, Tree is the default review view
      v0.27.16 | Themed diagrams on Mermaid 12, comment on any node or edge, patch-file review, embedded HTML documents render
      v0.27.15 | Plannotator TUI and Herdr Annotate announcement, element context on pinpoints, HTML links open as linked documents, All files panel, Classic diff default
      v0.27.14 | Pi plan progress survives compaction, Codex threads across rollout files, WSL browser setting, Mod+E edit mode

      What's New in v0.27.25

      A fix release driven by reports from people using the last two releases: code review with a colorized git config, Bitbucket reviews on large pull requests, wide tables in Firefox, Ask AI on Windows, and a failing install. Six pull requests, two of them from community contributors.

      Code review works with color.diff = always

      If your git config set color.diff = always, a common line in dotfiles, plannotator review showed 0 files with no explanation. Git wrapped every line of the diff in terminal color codes even though Plannotator reads it through a pipe, and the diff parser skipped every line it could not recognize. Viewed-file marks also stopped saving.

      Plannotator now turns color off for every git command it runs during a review, in every runtime. Setting color.ui=never alone is not enough, because a more specific color.diff=always wins over it, so each color setting is turned off explicitly. jj gets --color=never, gh pr diff gets --color=never, and the review agents and Call Flow analysis run their own git with color off. A normal diff is byte for byte what it was before, so saved comments and viewed marks are unaffected.

      (#1662, #1663, closing #1661, reported by @sholsinger)

      Bitbucket review fixes from real use

      @Toparvion tested Bitbucket Cloud review on a large pull request with several reviewers, existing comments, and Guided Review, and reported three things:

      • "View on Bitbucket after submitting" now opens. The browser only lets a page open a tab within a few seconds of a click, and posting a large review takes longer, so the tab was blocked as a popup. Plannotator now opens the tab the moment you click submit, sends it to the pull request once the review posts, and closes it if posting fails. The Feedback Sent screen also has a plain link. This affected GitHub and GitLab too, and links now also open from the VS Code panel.
      • Your overall comment shows at the top. Bitbucket's activity feed lists the newest comment first, so the summary comment posted before the inline comments ended up at the bottom. It is now posted after them, so it lands on top. A retry after a partial failure still never posts it twice.
      • Long reviews in Claude Code. Claude Code stops a slash command's shell after about 30 minutes. That limit belongs to Claude Code, not Plannotator. The Claude Code guide now shows how to raise it with BASH_DEFAULT_TIMEOUT_MS, and the review skill tells Claude it can re-run the review in the background. Your comments come back either way, because drafts are saved.

      (#1653, from #1583, reported by @Toparvion)

      Wide tables no longer collapse in Firefox

      In Firefox, commenting inside a wide table while the action bar was pinned to the top squeezed the table into a narrow strip beside the bar. The bar was both a float and sticky, and Firefox laid the table out around where the bar sat once stuck. The bar is now sticky without being a float, and tables keep their width in Firefox, Chrome and Safari. Plan review had the same problem and is fixed too.

      (#1660, closing #1659, by @ruaridhw)

      Additional Changes

      • Ask AI dropdowns are readable in dark mode on Windows. The provider, model and reasoning-effort lists showed light text on a light background. They now use the theme's own colors (#1658, by @Aayushyaash).
      • The install script reads the latest version correctly. It found the version with a line-based text match, which picked up the wrong value when GitHub returned its answer on a single line, and the install failed with a 404. It now reads the version by name, and the Windows install.cmd parses it as JSON. This fix has been live on plannotator.ai since it merged (#1656, closing #1655, reported by @jdppettit).

      Install / Update

      macOS / Linux:

      curl -fsSL https://plannotator.ai/install.sh | bash
      

      Windows:

      irm https://plannotator.ai/install.ps1 | iex
      

      Claude Code Plugin: Run /plugin in Claude Code, find plannotator , and click "Update now".

      Pi: Update @plannotator/pi-extension to 0.27.25 and restart Pi.

      OpenCode: Clear cache and restart:

      rm -rf ~/.bun/install/cache/@plannotator
      

      What's Changed

      New Contributors

      Contributors

      @ruaridhw tracked the Firefox table collapse down to a sticky float, measured it across browsers, and sent the fix.

      @Aayushyaash fixed the unreadable Ask AI dropdowns on Windows dark mode, with before and after screenshots.

      Community

      • @sholsinger diagnosed the 0-files review down to colorized git output and suggested the fix in #1661
      • @Toparvion tested Bitbucket review on a real team pull request and reported the popup, comment order, and timeout issues in #1583
      • @jdppettit reported the failing install, with a workaround, in #1655

      Full Changelog : v0.27.24...v0.27.25

    14. 🔗 r/LocalLLaMA Pi 1.0 released - MCP support now included by default rss
    15. 🔗 Rust Blog Demoting i686 Windows targets to std-only rss

      With Rust 1.100.0, the following changes to 32-bit Windows targets will happen:

      • i686-pc-windows-msvc Tier 1 with host tools target will be demoted to Tier 1 without host tools.
      • i686-pc-windows-gnu Tier 2 with host tools target will be demoted to Tier 2 without host tools.

      Builds of the standard library will continue to be distributed, but host tools such as the compiler will be no longer available. i686-pc-windows-msvc as a Tier 1 target still undergoes CI testing.

      To build 32-bit Windows binaries, cross-compiling from a still-supported host toolchain (such as a 64-bit Windows ones) will be required from now on.

      Background

      Desktop and Server 32-bit only x86 CPUs are no longer sold for over 15 years, and general 32-bit Windows support has ended in October 2025. This means that the development platforms these targets are meant for hardly exist these days, and even if they do exist they typically aren't capable enough for development.

      Even on the modern x86_64 hardware, building i686 Windows toolchains has proven to be problematic. We have encountered compiler binaries crashing when built with the i686 MSVC target, and the GNU C++ toolchain failing with OOMs during LLVM build.

      Considering all these things, cross-compiling these targets from a better supported one is what we have found to be the best solution forward. As part of that, we stopped producing host tools for these targets. For the time being, the prebuilt standard library is still available, and in case of i686-pc-windows-msvc still tested on CI.

      What Changes?

      After Rust 1.100, it will no longer be possible to install toolchains on 32-bit Windows hosts. We recommend cross-compiling from a still-supported host (such as a 64-bit Windows toolchain) instead. Other 32-bit platforms are not impacted by this change.

      For more details about these demotions, see RFC 3999 for i686-pc-windows-msvc demotion, and MCP 1020 for i686-pc- windows-gnu demotion.

    16. 🔗 New Music Releases Epica - Live at the Symphonic Synergy rss

      Epica - a new release is available:

      • 2026-10-02: Live at the Symphonic Synergy (Live)

      Amazon: Canada | Deutschland | France | United Kingdom | United States

      Visit muspy for more information.

    17. 🔗 New Music Releases Imminence - Axis Mundi rss

      Imminence - a new release is available:

      • 2026-10-02: Axis Mundi (Album)

      Amazon: Canada | Deutschland | France | United Kingdom | United States

      Visit muspy for more information.

  4. October 01, 2026
    1. 🔗 earendil-works/pi v1.0.0 release

      New Features

      • Fullscreen by default — The TUI now runs fullscreen. Set tuiMode to "regular" to keep the terminal's normal scrollback. See Terminal and display.
      • Leaner codemode — About 40% fewer prompt tokens, and errors that tell the model how to recover. See Codemode.
      • Image generation in codemode — Scripts call models.generateImages() with the session's credentials. See Generate images and Use image models.
      • Radius in/login — Sign in with Radius and set up its MCP server in one step. See Radius.
      • Anthropic copy code login — Sign in when the browser runs on another machine. See Authenticate interactively.
      • MCP OAuth hardening — oauth.authServerMetadataUrl, RFC 9207 iss checks, credentials per server, and step-up sign-in that keeps granted scopes. See Authenticate with OAuth.
      • Header-only quiet startup — quietStartup: "header" keeps the version and key hints and hides the rest. See Terminal and display.

      Added

      • Added an oauth.authServerMetadataUrl setting for MCP servers that advertise a wrong OAuth authorization server or none. Pi uses the configured metadata document instead of discovery (#10172).
      • Added quietStartup: "header", which keeps the startup header with version and key hints but hides the model scope line and loaded-resource listing.
      • Added models.generateImages() to codemode scripts. It runs image models such as OpenRouter's with the session's credentials and returns base64 image blocks that image() attaches to the result; usage counts toward the session cost like models.classify(). Extensions can call ctx.modelRegistry.generateImages(). See Use image models.
      • Added a copy code login method to Anthropic /login for headless setups where the browser runs on another machine (#10194 by @lucasmeijer).

      Changed

      • Changed the default TUI mode to fullscreen. Set tuiMode to "regular" or pass --tui-mode regular to keep the terminal's normal scrollback.
      • /login now offers "Sign in with Radius" at the top level, as the last option, with its status. After a Radius sign-in, /login offers to configure the Radius MCP server in the global mcp.json with "auth": { "provider": "radius" } and reloads. Cancelling a login returns to the menu it was started from.
      • The provider docs page is renamed to Providers, its "Cloud Providers" section is now "Provider Specific Config", and it documents Radius first.
      • MCP OAuth credentials are now stored per server name and URL, so MCP servers with the same URL can sign in with different accounts. Credentials stored by URL alone move to the first server that uses them (#10252).
      • Codemode costs far fewer prompt tokens: with the default tools and codemode active, a GPT-5.6 request shrinks from about 5,300 to 3,300 tokens. The codemode description lists the script globals in one line each and points to the new Codemode reference for the models API, which the model reads when it needs it. Declared tools say in one line how scripts call them and what the call resolves to, instead of repeating their full declaration, and the system prompt's codemode guidance and MCP server section are shorter.
      • Codemode errors now say how to recover: reading a tool or models member that does not exist names the close matches (tools.Bash suggests tools.bash), models.classify() and models.generateImages() reject malformed arguments with the expected shape, an unknown model points to models.getAvailableOfType(), an oversized store() value explains what the store is for, and a script that generates images without showing them gets a note. Scripts that probed for a tool with typeof tools.name must use "name" in tools.
      • /login and /logout now label providers without credentials as "not configured" instead of "unconfigured".
      • OAuth browser pages now show the color Pi logo.

      Fixed

      • Fixed MCP OAuth sign-in accepting an authorization response whose iss parameter names another authorization server; the code is now rejected before it is exchanged (RFC 9207).
      • Fixed MCP OAuth sign-in failing with Invalid scope when the token response contains "scope": "", and similar failures for other empty or null optional OAuth fields (#10266).
      • Fixed the sign-in URL printed by /mcp login not being clickable when it wraps (#10186).
      • Fixed --provider without --model being silently ignored and running the default model from another provider; it now fails with an error (#10236).
      • Fixed MCP servers that ask for more scope (insufficient_scope) requesting sign-in over and over. The new sign-in requested only the missing scopes, so the new token lost access the previous one had; it now keeps the granted scopes.
      • Fixed user messages in the transcript keeping two full-width copies of every rendered line; they keep one, with identical output.
      • Fixed /login and /logout labeling every OAuth sign-in, including Radius, as a subscription; only subscription-backed providers say "subscription", other OAuth sign-ins say "account".
      • Fixed the startup header logo rendering with gaps in Apple Terminal; it now shows a colored "Pi" with the version instead.
      • Fixed deferred MCP tools that tool_search loaded being dropped on resume and /reload even when their server reconnected before the next prompt, because the session restored its tools before the MCP servers reconnected.
      • Fixed the system theme making pastel palettes such as Catppuccin Frappe much more vivid; palette colors now keep their chroma (#10255, #10293 by @dgtlntv).
      • Fixed slash command autocompletion not triggering when the input starts with whitespace (#10218 by @haoqixu).
      • Fixed color bleeding past mouse selections and search highlights in fullscreen mode when a styled token ends at the highlight boundary (#10169).
      • Fixed memory retained per rendered message in the transcript; a long assistant message keeps about a fifth of the heap it kept before.
    2. 🔗 anthropics/claude-code v2.1.287 release

      What's changed

      • Added Claude Mods: plugins may now modify deeper behavior
      • Added You should know, a built-in mod where a side agent watches your back and flags things you or Claude might miss. Turn it on with /plugin enable cc-plugin-you-should-know@builtin (for first-party sessions with telemetry on)
      • Added an n:<text> filter to the agents view that matches session names and tasks; a filter now shows matches in collapsed sections and Enter opens the first match
      • Added prompt_text to the OpenTelemetry user_prompt event, a copy of prompt for backends that nest dotted keys; drop or mask it wherever you drop or mask prompt (#70763)
      • Added URL prompts from MCP servers on the 2025-11-25 protocol, for example to sign in. If a server no longer connects after this update, add "bareElicitationCapability": true to its MCP config entry
      • Windows: Added a startup warning when denying the Bash tool also turns off the PowerShell tool, so Claude has no shell tool
      • Self-hosted runner: Added a built-in gh api (REST only) for sessions that use Anthropic-managed git on macOS and Linux machines where the GitHub CLI is not installed
      • Fixed fast mode staying off in remote sessions owned by an agent with no user account, even when the organization allows it
      • Fixed Remote Control not receiving messages for minutes at a time when a reconnect request got no response; it now gives up after 30 seconds and retries
      • Fixed hooks configured with asyncRewake waking Claude over and over with "found issues" notifications when the hook's script file is missing; the broken hook is now reported once
      • Fixed tool heartbeats not reaching SDK hosts while the model's response stream was stalled with no data arriving
      • Fixed Bedrock and Vertex startup model checks ignoring an enforced availableModels list, which could collapse /model to one Opus row
      • Fixed the Claude in Chrome browser picker showing a JSON parse error when Chrome could not be reached
      • Fixed picking Fable in /model on a claude.ai login saving the current version's id, so your saved default now follows the newest Fable like Opus and Sonnet do
      • Fixed switching between Opus 5.5 and Sonnet 5.5 (/model, opusplan) rewriting earlier MCP tool announcements, which could drop earlier extended thinking
      • Fixed Amazon Bedrock Guardrails blocks that arrive mid-response ending the turn with an API error instead of the guardrail's message when the reply began with thinking
      • Fixed a dangerous rm (such as one on / or the home directory) losing its always-ask safeguard when the same command also redirected output to a ~ or wildcard path
      • Fixed claude -p and SDK sessions repeating a model fallback on every later message after the model was switched while a reply was running
      • Fixed a folder's CLAUDE.md being attached a second time after resuming a session or after a compaction
      • Fixed background sessions that could not be reopened from claude agents after the agent exited and removed the worktree the session was started in
      • Fixed /advisor pairing checks: Sonnet 5.5 can now advise Opus 4.7 and 4.8, and advisors the API would refuse are flagged up front instead of being silently dropped
      • Fixed Bash permission prompts showing internal parser names such as "Contains simple_expansion" instead of a plain explanation
      • Fixed a cause of fullscreen sessions on slow or busy machines exiting with "Claude Code exited after an unrecoverable interface error" while a scroll key was held in a long conversation
      • Fixed organization per-tool permission ceilings being silently dropped for an MCP tool named __proto__
      • Fixed Claude being told to page large MCP results saved as JSON with Read's offset and limit, which cannot split one long line
      • Fixed the commit attribution reminder being delivered inside a tool result after a compaction
      • Fixed screen reader mode leaving the cursor away from the typed text in search boxes (such as /resume and /permissions) and sign-in code fields
      • Fixed screen reader mode refusing Enter with nothing typed on /rewind's summarize options, whose added context is optional
      • Fixed screen reader mode showing a "Tab to amend" hint on approval prompts, where Tab does nothing
      • Fixed screen reader mode listing arrow keys that do nothing in /permissions and /mcp, and saying "Select with numbers" in empty menus or while a search box has the keys
      • Fixed screen reader mode leaving out the changed lines in file edit approval prompts and other diffs
      • Fixed screen reader mode sending the claude --teleport progress screen, and an MCP form field while it is being checked, to the screen reader again on every spinner frame
      • Fixed CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS not removing the structured-output format from session-title and prompt-hook requests, which Bedrock-backed gateways reject
      • Fixed screen reader mode leaving out the top lines of a second approval prompt, a changed /config row or the rejected-plan line when the previous screen was taller than the terminal window
      • Fixed --include-partial-messages sending a cut-short reply's message_stop late or never, so apps could show the reply as still in progress
      • Fixed claude agents sometimes not showing the permission prompt a background session is waiting on
      • Fixed /ultrareview giving advice about .git/info/attributes when the upload stops on a committed .gitattributes it cannot read, such as one saved as UTF-16
      • Fixed claude remote-control failing to register behind an HTTP proxy with a misleading "Check your organization permissions" error (#97352)
      • Fixed sandboxed Bash commands on Linux inheriting an open handle on the Claude Code executable
      • Fixed the running-tool dot and three spinners still moving with the "Reduce motion" setting on, and /rewind's confirm screen updating its "ago" time while you type a note
      • Fixed times in claude agents changing every second in screen reader mode; they now change at most every 10 seconds
      • Fixed a revoked claude.ai login showing a generic API Error: 401 instead of "OAuth token revoked"; in -p mode the error now starts with "Failed to authenticate"
      • Fixed /ultrareview upload refusals advising you to copy a variable named by a repository's settings file into your own user settings
      • Fixed --output-format stream-json and the SDK not streaming the turns of a context: fork skill run by typing /<skill> as the prompt, as they do for the Skill tool's fork
      • Fixed /feedback and /bug: the pre-filled GitHub issue no longer includes your recent error messages, and the confirmation screen now lists them as part of the report
      • Fixed claude plugin marketplace add --sparse and git-subdir plugin installs failing with "transport 'http' not allowed" when the repository is served over plain http
      • Fixed cloud sessions sometimes losing the earlier conversation when the session restarted while it was being compacted
      • Fixed a plugin reload that overlapped the startup --plugin-url download corrupting the session's cached copy of the plugin archive
      • Fixed /desktop quoting partial output when opening Claude Desktop timed out or printed too much output; the error now names the cause
      • Fixed an MCP connector tool call occasionally running twice, or the connector's calls failing until restart, when its server changed which MCP protocol version it supports
      • Fixed SessionStart hooks from synced plugins not running in new cloud sessions
      • Fixed the transcript's "N hooks ran" summary and the verbose debug log's matched-hooks count including Claude Code's internal callbacks, so one configured hook no longer shows as two
      • Fixed files Claude sends from cloud and Remote Control sessions failing when the upload finished just after the 30-second timeout; it now waits 35 seconds
      • Fixed repositories added mid-session in cloud and SDK sessions not loading their skills and plugins, and loading CLAUDE.md late, after Claude changed directory
      • Fixed PNG, JPEG and WebP images over 8,000 pixels on a side failing to send from a remote session; Claude now sends a scaled-down copy
      • Fixed messages sent from the Claude apps with 17 to 20 attached files delivering only the first 16
      • Fixed headless sessions reporting an MCP server as needing authentication after one refused call, even though later calls succeed
      • macOS: Fixed Remote Control sessions started with claude remote-control stopping mid-turn when the Mac went to idle sleep
      • Windows: Fixed interactive claude hanging or crashing with "Raw mode is not supported" when its input is piped or redirected; it now says why and exits (use -p for piped input)
      • Bedrock, Vertex, Mantle: Fixed model availability checks under CLAUDE_CODE_SKIP_*_AUTH sending a different Authorization header than real requests when ANTHROPIC_CUSTOM_HEADERS repeats it
      • Improved /config: settings that cycle show ‹ › and step both ways with ←/→, narrow terminals stack each value under its label, and PgUp/PgDn page the list
      • Improved plugin marketplace errors to say in plain words why a marketplace was ignored or refused, and what to do
      • Improved plugin listings to note when a plugin's dependencies were not installed, and updating a plugin now retries an install that did not finish
      • Improved the Claude apps gateway's error when Amazon Bedrock rejects a model ID: developers now see which model is unavailable, and the gateway log names the ID that was sent
      • Improved SDK sessions so a message sent with priority "now" no longer cancels a running web fetch or web search; it keeps loading in the background
      • Improved /memory: the left and right arrow keys now flip its on/off settings, such as Auto-memory
      • Improved /skill names typed mid-message: Claude is now told they are skills, including disable-model-invocation ones
      • Improved the contrast of the prompt input border in light themes and of the ❯ before your earlier messages
      • Improved delivery of files Claude sends from cloud sessions and Remote Control: an upload that fails on a timeout, a network error or a 502, 503 or 504 is now retried once
      • Improved what Claude says when a file cannot be sent for a reason that may be temporary: it now mentions that you can ask for the file again in a few minutes
      • Improved the prompt for a held message from another session to show the message between dashed lines, matching other permission prompts
      • Improved MCP and other tool permission prompts to show the tool call between dashed lines, matching file edit prompts
      • Improved MCP startup in headless mode: a remote server whose first connect fails transiently is now retried without waiting for the slowest server to finish connecting
      • Improved files Claude sends from a remote session: large files now stream from disk instead of being read into memory, and a file over the size limit is refused with the server's limit named
      • Improved the explanation Claude gives when the server refuses a file it sends from a Remote Control or cloud session, such as an oversized image
      • Improved handling of large MCP tool results: less memory, smaller session files, and no extra upload to count tokens for results far over the limit
      • Windows: Improved Bash tool speed by removing a subshell that ran before every command
      • Changed a shell write through a repo-committed symlink onto a sensitive file or out of the working tree to name where it lands and wait for a person, on lines with a ~ target too
      • Changed Opus 4.7+ and Fable to use a 1M context window by default on Bedrock, Vertex, Foundry and the Claude apps gateway, with no [1m] suffix (CLAUDE_CODE_DISABLE_1M_CONTEXT=1 keeps 200K)
      • Changed replies from claude agents to arrive as queued messages; slash commands other than /stop sent while a turn is running now run when it ends
      • Changed whole-tool Bash allow rules and allowing hooks to prompt for, not run, shell writes to files Claude Code's file tools refuse outright (the Anthropic profile store, the host credentials file)
      • Changed right-click paste on Windows and Linux, and middle-click paste on Linux, to happen when the button is released; moving the pointer away before releasing cancels it
      • Changed MCP server alwaysLoad: false to defer all of that server's tools behind tool search
      • Changed screen reader mode to write new or changed lines without first pausing with the cursor at the start of the line; set CLAUDE_AX_PREPARK_MS=50 to restore the pause
      • Changed automatic model switches after a flagged message to keep your current effort level instead of the new model's default
      • Changed waiting permission prompts to show oldest first, so a new prompt no longer covers the one you're reading (prompts with a countdown still open on top)
      • [VSCode] Added "Run in background" to a running command or sub-agent, to move it to the background and keep working
      • [VSCode] Added the output of background shells and Monitors to their cards in the agent map
      • [VSCode] Fixed settings dialogs blaming a timeout when Claude Code's reply was too large to confirm a save
      • [VSCode] Fixed reopening a cloud session that the side bar already brought to this machine opening it again in a new tab; the side bar is shown instead
      • [VSCode] Fixed the side bar's Web tab not listing cloud sessions started after the window loaded; a failed load now says "Remote server is not connected" instead of "No web sessions yet"
      • [VSCode] Fixed a tab restored after a reload starting a second Claude process on a conversation the side bar already has open; it now shows the "still open somewhere else" notice
      • [VSCode] Fixed tool-row file links, session-list links and two hints showing in plain text
      • [VSCode] Fixed a background agent's still-running command showing as failed once the main turn ended
      • [VSCode] Fixed a user's own /usage or /context command opening the extension's dialog instead of running when picked from the command menu
      • [VSCode] Fixed file links in the plan preview tab doing nothing when clicked; they now open the file like links in chat replies
      • [VSCode] Fixed opening a tool's input or output in an editor tab failing with "Timeout waiting after 1000ms" on remote hosts such as WSL when the tab is slow to appear
      • [VSCode] Improved the Manage plugins dialog: a failed marketplace add, remove or refresh now says what went wrong
      • [VSCode] Changed the Claude in Chrome "Enabled by default" switch to also connect the editor's own sessions, which still ask before browser actions
      • [Cloud sessions] Fixed occasional failures to fetch from or push to GitHub when GitHub briefly refused a newly issued access token
      • [Claude Tag] Fixed Claude posting a failure warning, such as a spend limit notice, in a Slack thread when a background event like GitHub activity woke it and nobody was waiting on a reply
      • [Claude Tag] Fixed Claude Tag's spend limits page in admin settings leaving out recently created and private channels in organizations with many channels
      • [Claude Tag] Improved Claude's task list in long Slack threads: background work no longer reposts it as a new message on its own, so people following the thread aren't notified
      • [Code Review] Fixed finding comments and their "Why this was flagged" text stopping mid-sentence; they now end on a complete sentence
      • [Code Review] Fixed Code Review skipping a pull request after a new push when its review had failed twice on the previous commit; it now reviews the latest commit
      • [Code Review] Improved the failed-review card on a pull request whose conversation is locked: it now says the lock blocked the review and that nothing was posted or charged
    3. 🔗 r/LocalLLaMA Clef: Open Weights decision model by Cloudflare rss

      Clef: Open Weights decision model by Cloudflare | submitted by /u/paf1138
      [link] [comments]
      ---|---

    4. 🔗 r/LocalLLaMA Heretic is on PewDiePie! rss

      So I haven’t played a computer game in 20 years, and I know nothing about Minecraft, and I definitely prefer classical literature over YouTube culture, but even I have heard about the individual called PewDiePie, for two reasons:

      1. His monicker starts with the initials of my own name

      2. I remember a recurring Internet meme a few years ago where he was competing for the most subscribers with an Indian film music channel

      I had never watched a single one of his videos, however.

      Well, until today, when people started spamming me with messages informing me that Mr. Kjellberg aka PewDiePie has tried outHeretic and made a video where he talks about it:

      https://m.youtube.com/watch?v=ODDJXGY_1kQ

      (Heretic mentioned around 9:00)

      Obviously I’m thrilled that a less technical audience is being exposed to my work, and the more people understand what is possible the better. I expect to be receiving a couple hundred more mails in the coming days asking how to run Heretic on ChatGPT (you can’t) , or accusing me of working for the CIA (I don’t) , but other than that, the more the merrier I guess 😏

      Heretic 2.0 coming soon…

      submitted by /u/-p-e-w-
      [link] [comments]

    5. 🔗 The Pragmatic Engineer The Pulse: RoR creator sparks new “death of coding by hand” debate rss

      Before we start: if you happen to be in San Francisco on Thursday, 5 November, join me on the System Update with The Pragmatic Engineer event. This is an evening with OpenAI, Linear and DoorDash and myself, organized by Sentry. We get into what AI-augmented automations they're running in prod, and how it's going, in an off-the record (that is: not recorded!) and raw conversation. Seats are limited, and you can RSVP here.


      Hi, this is Gergely with a bonus, free issue of the Pragmatic Engineer Newsletter. In every issue, I cover Big Tech and startups through the lens of senior engineers and engineering leaders. Today, we cover one out of four topics from the last week 's issue of The Pulse . Full subscribers received the article below seven days ago. If you 've been forwarded this email, you can subscribe here .

      The creator of Ruby on Rails, David Heinemeier Hansson, caused quite a stir last week with comments in his Rails World keynote, when he revealed that coding by hand is dead at his company, 37signals.

      This is a big deal because 37signals created Ruby on Rails, and they are known for their software craft there, especially when it comes to code quality. It's also a business that's 27 years old and is profitable. Despite that pedigree, DHH caused a stir among the dev community, saying:

      "At 37signals, a couple of weeks ago, we made the decision that it clearly means we 're done writing code by hand. We have gone pencils down on the idea that we were gonna write code by hand, as a normal course of business creating things.

      Writing code by hand at 37signals is now an exceptional state. It is like seeing a bug in Sentry: something here went wrong; why was the agent not able to produce what we wanted? Okay, maybe for a little while, we'll still get the old pencil out and dot it down for them, but then we fix the machine, we fix the factory, we get things going again. This is a recognition of what's already happening."

      DHH compared the maturation of AI tools into being highly capable at coding with the impact upon the craft of painting of the arrival of the camera:

      "On November 24th, 2025, we got the "Kodak Brownie" of our era. We got Opus 4.5. AI technology, accessible in a harness that many people could afford to use and experience for the first time what it's like to create software in pairing with a new form of intelligence. This was the tipping point for me. There was everything before November 24th, and then there was everything after. This is going to be the date that history books going forward will mark as the inflection point for the age of agents."

      He shared how 37signals has embraced a future where coding by hand is almost entirely absent:

      • Embracing native mobile apps instead of web: famously, 37signals is bearish on native iOS and Android apps and has built web versions instead. With AI, they are betting on native apps being much easier to be built with a small team and are already building new ones.
      • Moving backend services to Rust, not Ruby : This is due to performance reasons and because agents write good enough Rust. That's remarkable to hear from the creator of Ruby on Rails!
      • Ruby on Rails remains for web apps : 37signals is not leaving RoR behind, but only because Ruby on Rails' convention-over-configuration design makes it easy for agents to work with it.

      DHH closed by revealing that he no longer even thinks of himself as a professional programmer (emphasis mine):

      "I have retired from being a professional programmer. I think it was somewhere around 4 to 5 months ago, maybe March. I spent a quarter of a damn century chiseling code by hand and loving every moment of it. This is not something to look back upon with regret; this is something to look back upon with joy and accept that it is over.

      Writing code by hand is no longer an economically productive enterprise for the vast majority of programmers working at the vast majority of companies. On the other side of that is a new career as a professional maker of things, steering intelligence that was only available in science fiction up until a few moments ago.

      One of the things we're gonna have to revisit is everything we think we know about software architecture. The main tool that we've used for a very long time is abstractions. Abstractions don't make quite the same sense in the age of agents. The reason we did abstractions was in part not to repeat ourselves; well, now the price of repetition has gone to near zero."

      It's worth noting DHH's keynote chose a spicy topic for a conference attended by engineers who are personally and professionally invested in the craft of building software!

      Decline of coding by hand is long predicted

      In the first issue in The Pragmatic Engineer this year, on 6 January, I wrote:

      "When AI writes almost all code, what happens to software engineering? No longer a hypothetical question, this is a mega-trend set to hit the tech industry. (...)

      The bad news is that change will probably be rapid. It's barely been a year since the idea of Claude Code was born in Boris Cherny's head, and already similar tools like OpenCode, Codex, Factory, Amp, Cursor, and more capable agents are changing how software is written. Change has always been part of working in tech, but I cannot recall it being this fast, or happening across the whole industry at once!"

      I concluded that this change was on its way, based on my own experience of building software with Opus-4.6 and GPT-5.2, and from talking with experienced engineers who had resisted "AI hype" for good reason, but who had come to see that AI can now generate code that's "good enough" in many cases.

      Back then, I made a few predictions about what will happen when AI agents are producing most of the code for engineers:

      • Sloppier code
      • Weak software engineering practices hurting sooner
      • "Coders" who are not software engineers see less demand
      • Tougher work-life balance for engineers
      • Junior engineers pushed to become seniors, fast
      • Computer science education increasingly required for new hires
      • A massive explosion in code and software, for which someone must be accountable

      So far, it's a messy transition and we engineers are responsible and accountable for a lot more code that we didn't write, but which is in production anyway.

      Non-engineers also getting into agents

      At the end of January, I shared a deepdive that was pretty close to home for me: my brother's 30-person, 15-engineer startup, Craft Docs, made its own sharp pivot to AI by building their own AI harness for non-engineers - called Craft Agents - two weeks before Claude Cowork was released, and months before ChatGPT Work launched.

      Craft resisted the temptation to use AI when it did not feel productive, but with the model releases of November 2025, they found LLMs are not only useful for coding, but also for non-engineering work like customer support. In the deepdive, I went into more detail about non-engineering use cases (which engineers enabled) like:

      • Automatic triaging of bug reports with agents
      • Data enrichments added to all workflows
      • Customer support "skills" like processing feature requests
      • The marketing team building websites without devs
      • HR automating tedious work
      • Finance automating personal workflows

      Craft Docs seemed early to a trend that has become more widespread, by having both their own engineering and non-engineering folks onboard to an AI harness. Now, there are signs other companies are doing the same: at OpenAI, non- engineering units like finance, recruitment, and legal moved over to Codex in June 2026:

      altNon- engineering teams moved over to use OpenAI 's AI harness, Codex (now renamed to ChatGPT Work.) Source: Inside OpenAI 's software factory

      In some ways, it could be comforting to know that it's not only software engineering where the tools and workflows are quickly changing: every other function in tech is experiencing the same!

      It 's messy right now

      Just last weekend, a rant by an anonymous engineer in Big Tech hit a nerve with many people in the industry. An engineer with the username voxium posted (emphasis mine):

      "The state of engineering right now is horrible. It has been half a month since I started a new role at a big company.

      Nobody knows anything here. The specs, code, tests, PRDs, tickets, resolution of those tickets, reports, etc., everything is made by Claude Code. Nobody on my team likes this.

      They are being forced to ship as much as they can. I have heard multiple times from higher management that pushing code is not a bottleneck, so why are we slow?

      People are working 12 to 13 hours a day just to press enter. Nobody is reading anything. Humans in corporate are doing nothing on their own.

      Everyone, literally everyone, from an L1 to an L7 engineer here is doing the same thing. Talk to Claude.

      There is no sense of victory. Nobody is resolving bugs. In reality, nobody is thinking anymore. Everything is done by LLMs. It is so soul-sucking.

      I would not mind it, to be honest, if we were at least given the time to check out the code and see what is going where. But no, the goal is to just ship. No matter what happens."

      This post rings true because it is happening at many places where there's more AI usage, engineers do "outsource" thinking to LLMs, and end up not caring about anything else except shipping something to production.

      Quality in decline

      Since the beginning of the year, the quality of software has been degrading pretty much everywhere, much of it caused by over-reliance on AI, or perhaps more accurately, the outsourcing of thinking and decision making to AI. In July, I moved my video podcast off of Spotify after a series of unexplainable outages, and Spotify's engineering team seemed to take no real pride or accountability in fixing the root causes of the issue.

      Only this week, Uber shipped a new feature to production in the Uber Eats app - a new way to select extras with your food order - with seemingly no QA testing:

      altUber Eats this week, when I attempting to order a burger. Can you spot two obvious, sloppy bugs on this page?

      Inside this new "add-ons selector" in Uber Eats, I noticed three bugs at once:

      • " Choose up to 999": no engineer, designer, or PM bothered to check what happens when a restaurant does not fill out a number on how many toppings to add, or adds a ridiculously large number. The most toppings that my screen allowed to be selected was six and not 999, so I could not even choose the option of 999 buns for my burger.
      • Sloppy overflow. A rule of thumb, during my time at Uber, was that text will never overflow, even when localized. Basilcummaynaise (basil mayo in Dutch) broke this rule but still shipped.
      • Functional bugs in the selector. I originally tried to order from my favorite Mexican place: a bowl with no rice or bulgur as the base. There's the option to select "rice", "bulgur" or "nothing" as the base, but selecting "nothing" counts as an extra side, and the app doesn't allow the ordering of a bowl with no base.

      I've used the Uber Eats app for years, and this was the first time I saw such a sloppy feature release. I assume that devs and PMs building it have all "checked out", stopped doing proper QA, and assume that the agent will take care of all of it. There's no other way to explain three bugs shipped to all customers but seemingly noticed by nobody until I posted about it. To the Uber Eats team 's credit, they reached out and are looking into fixing all three issues.

      Software engineering to be more important than ever

      I'm personally past the shock and grief stages of agents taking over the activity of coding. At first, I assumed this change would reduce the amount of work for engineers. But, counter-intuitively, that actually seems to be growing:

      • We need to understand the characteristics of LLMs better. LLMs feel familiar as they can produce text in a way only humans could do before. But they are less reliable, still prone to hallucination, suffer from capability gaslighting, and many other problems. They can also be expensive and slow.
      • New systems need to be engineered. Agentic "software factories" can now produce code, based on the input provided. But how is this code validated? How much can detecting defects or various issues be automated? This is a brand new area, and we need to build new types of systems, often based on old ideas. One such example is OpenAI's software factory, the other one is Ramp's Inspect internal coding harness.
      • Nondeterministic LLMs can generate deterministic code. An area I feel is under-explored and under-appreciated is the use of LLMs to substitute LLMs usage in agentic "software factories" with deterministic code they generate. For example: instead of running AI code review that is expensive and slow on all PR requests, could AI generate linters that catch the majority of common issues? If this is possible, complex lint rules would run faster, be more reliable and cheaper to execute than LLM calls.

      New categories of systems and products will be built by engineers who "get" LLMs and AI engineering. We are seeing the majority of venture funding pour into AI companies because AI creates new business models, new revenue streams, and disrupts "traditional" software. For example, who would have thought that companies would spend tens of thousands of dollars, per engineer, on AI coding tools? Or that the category of AI inference providers would become as massive as it already is from barely existing two years ago?

      This technological change will re-jig parts of the tech industry: the winners will surely win big, and teams and companies choosing inaction could be out- executed and displaced by nimble competitors. And in many ways, this is great news for us software engineers who keep up with the technology. Companies are now investing in innovation and are willing to pay top-of-market for software engineers who can help them build AI products or become AI-native.


      Read the full issue of The Pulse this is from, or check out this week 's The Pulse. This week's issue covers:

      1. Firebase: global outage & poor handling by Google. The Firebase iOS SDK crashed after a backend change, crashing all apps which used Firebase analytics for 2-6 hours. Google did not update the status page or offer any postmortem, which is a head-scratcher from a company known for standout incident management practices.
      2. OpenAI 's platform play from AWS playbook? OpenAI is becoming a platform where it's possible to allocate ChatGPT spend on open models and AI offerings from among 16 partners, not just OpenAI models. It's fair to ask if Anthropic will consider a similar platform play.
      3. More data on companies moving to open models. Vercel's AI gateway shows 60% of model spend goes to open weight models, and OpenRouter also shows open models are being more used than closed ones.
      4. Why CTOs and VPEs are quitting en masse: another take. What if it's not "founder mode", but about people who love building software, feeling like they can do it solo (or with a small team) with AI tools?
    6. 🔗 @HexRaysSA@infosec.exchange The IDA 9.5 Beta is live! mastodon

      The IDA 9.5 Beta is live!
      ◾ IDA MCP + Assist for agentic RE
      ◾ New TriCore, Hexagon and DEX/ODEX decompilers
      ◾ Recursive decompilation
      ◾ New Malware Analysis add-on
      ◾ And more...

      Beta members: it's in your Download Center now.

      Not in the Beta program? Join from the customer portal today.

      https://hex-rays.com/blog/ida-9.5-beta-is-available

    7. 🔗 Hex-Rays Blog IDA 9.5 Beta is available rss

      IDA 9.5 beta

      IDA 9.5 Beta is now available

      The IDA 9.5 Beta is live, starting today. If you're part of our Beta Program, the new build is already waiting in the Download Centerof your customer portal, so you can put it to work right now. Not enrolled yet? You can join the program in a few clicks from your customer portal dashboard, and your testing is what tells us a feature is ready for production, where a regression slipped in, and what still needs sharpening before release.

    8. 🔗 HexRaysSA/ida-nexus v0.13.2 release

      What's Changed

      New Contributors

      Full Changelog : v0.13.1...v0.13.2

    9. 🔗 exe.dev ETOOMANYTHINGS? Run Fewer Agents rss

      At a conference last week, I sat through a bunch of software factory demos, and they were all about task management. Kanban boards, Slack interfaces, email interfaces, dependency graphs, ticketing, graphical interfaces, you name it.

      Why is everyone building task management and calling it a software factory? Agents are slow to execute. And the obvious, easy fix to latency is to hide it by starting a new agent every time you get blocked.

      But concurrency is really rough on humans. It’s stressful. It trashes flow state and thrashes our mental page caches.

      Task management isn't a solution; it's a band-aid.

      This is a call to arms. Let’s make it possible to be equally productive with fewer agents.

      Bitter Lesson to the Rescue?

      As models improve, “good enough” models will get ever faster. This will naturally reduce latency, and in turn concurrency. We’ll still have to do some old-fashioned engineering, like making tests run fast.

      But we don't have to wait! Here are some things we’ve experimented with at exe.

      Use a Fast Model for Talking With Humans

      Start tasks with a powerful model. But instead of reading a wall of text and writing an essay in response, put all the comms in the hands of a fast, competent model like Luna 6. Give Luna no coding tools, and make its context window intentionally short and focused on what the human conversation requires. Then, when you go to check in on a task, you can have real-time discussions with Luna. With a fast feedback cycle, attention doesn’t wander, and you can stay engaged for longer. Eventually, Luna exhausts its ready information, and the beefier model takes back over, with rich, substantial user feedback to work from.

      This works reasonably well. But we found that, shock of shocks, absorbing information via chat was rather constraining. It was frustrating to have communications be squeezed through a tiny pipe. Luna might be an exciting, bendy straw, but it’s still a straw.

      Screen Recording

      That led us to focus on the HCI aspects of the problem. Back when I coded by hand, I would stare intently at screens full of text, processing it slowly, navigating freely between files, thinking. But the UI that is presented by most coding harnesses is a single linearized text thread, typically interspersed with lots of irrelevant tool call noise. There’s very little human agency or control, and no organization. You can’t even rely on the most important content being at the bottom: noise from straggler subagents often drowns out the agent's primary response.

      So: How can we restore human agency and optimize for human I/O, rather than making the human adapt to the agent?

      For input, the human retina is a powerful information processing system, when given structure to work from. That is, not a wall of text. I can still pick a Go panic stacktrace out of terminal logs scrolling by at 30fps.

      For output, even for the fastest typists, typing is typically slower than speaking. Also, typing requires coordination with lots of other parts of the computer. The text input window has to be focused. Your cursor must be in the right place. It constrains the other things you can do with your keyboard and mouse and what you can look at. Many coding agents, including Shelley, worked around this with affordances for annotating text, diffs, and web pages, but it still requires clicking, and it fundamentally requires cooperation from everything else.

      We recently launched a simple show-and-tell feature in Shelley that is tailored to humans. Click the camera button. Shelley records audio so that you can think out loud as you explore, and it records your screen. Anything on screen is a thing that you can point at, talk about, refer to. Poke around, muse, backtrack, trail off, resume, change your mind, whatever. Agents are extraordinarily good at understanding rambling.

      When you stop recording, Shelley transcribes all of your audio, including word-level timestamps, and makes a contact sheet from the video. It correlates your words with what was on screen when you said it. The computer does the hard work of collating all of the feedback.

      In my experience, this form of interaction is typically deeper, richer, longer, and more engaged than chatting with an agent. This is particularly useful for UI work or anything visual.

      DOM Recording

      What if you're not working on a visual task? The same ideas can be adapted nicely to other forms of engineering work.

      Another feature of human cognition is that we work well with concrete examples, not abstract descriptions. Agents are happy to be told, but humans prefer to be shown. Humans also benefit from diagrams, charts, and visually structured diffs.

      With this in mind, I have been using a slightly different system that we haven't shipped in Shelley (yet). Instead of spewing all its output to me in raw text, the agent generates an HTML artifact. This HTML is standardized to reduce visual noise, ruthlessly reduce verbosity and complexity, structure information visually, and make navigation easy. Typical sections include notable decisions made, open questions, examples, and git diffs.

      This HTML artifact includes a record button. When clicked, it records audio. Instead of using screen recordings, it uses a cheaper, more precise in-page DOM tracker JavaScript that tracks what is being viewed, scroll position, where my pointer is (if there is one), what text is highlighted, where I click/tap, all with timestamps. As before, it correlates timestamps with audio and feeds it all back to the agent.

      I find that this reduces mental overhead; I am freed to focus on the content.

      More Latency Hiding

      Because really engaging with content takes time, there is an additional latency-hiding trick one can pull. I am still experimenting with this, but the idea is to take prefixes of my feedback while it is still being recorded, send them to the agent, and livestream its responses back to a tab on the very same HTML artifact I'm already looking at. By the time I’ve spent 10 minutes reading and probing, I will inevitably have asked questions and provided provisional guidance. Before I context switch away, some of my questions have answers waiting. Some of my design decisions have follow-up questions. And the camera is still rolling. This sneaks in an extra round trip with the agent without breaking my concentration.

      There’s Still Work to Do

      I am not yet at a point where I have the same net throughput with one or two agents that I do with many, but my attention is being restored to me, little by little.

    10. 🔗 smol-machines/smolvm smolvm v1.22.1 release

      What's Changed

      • Bump libkrunfw to the guest kernel with conntrack marks so nft ct mark rules load by @BinSquare in #1493
      • Stop the workload wait ticker without waiting out its sleep by @LoganGrasby in #1495
      • Pin libkrunfw to the merged conntrack-mark commit on main by @BinSquare in #1494
      • Restrict serve mTLS clients by certificate subject CN by @BinSquare in #1499
      • Speed up restored starts with --keep-identity and RAM prefetch by @BinSquare in #1497
      • Preserve HTTP machine records when storage cleanup fails by @hfiguera in #1498
      • Give machine update the --outbound-localhost-only flag create has by @BABTUNA in #1500
      • Prefetch restored RAM and allow keep-identity restores in the embedded runtime by @BinSquare in #1502
      • Bump the workspace to 1.22.1 by @BinSquare in #1496

      New Contributors

      Full Changelog : v1.22.0...v1.22.1

    11. 🔗 backnotprop/plannotator v0.27.24 release

      Follow @plannotator on X for updates

      Missed recent releases? Release | Highlights
      ---|---
      v0.27.23 | Bitbucket Cloud PR review, Question UI for answering agents in place, opt-in auto-update, viewed files remembered, review another repo or worktree
      v0.27.22 | Plans open the docs they link to, Claude review jobs locked down, Code Tour on Linux without Claude's sandbox, Pi reviews the latest plan
      v0.27.21 | Remote and phone sessions load several times faster, real Request changes on GitHub, model pickers show real names, OpenCode fixes
      v0.27.20 | Mistral Vibe support, annotate gets the full Options menu and Settings, jj Commits panel, long lines wrap in plan code blocks
      v0.27.19 | Before/After image previews in code review, file comments as GitHub file threads, forge-correct #123 links, /plannotator-last finds the right session
      v0.27.18 | Model pickers from your installed Claude and Codex (Opus 5.5, Fable 5.1, GPT-6), unsent PR review comments survive new pushes
      v0.27.17 | Diagram files open in the diagram viewer, OpenCode switches model with agent, idle review stops polling the git remote, Tree is the default review view
      v0.27.16 | Themed diagrams on Mermaid 12, comment on any node or edge, patch-file review, embedded HTML documents render
      v0.27.15 | Plannotator TUI and Herdr Annotate announcement, element context on pinpoints, HTML links open as linked documents, All files panel, Classic diff default
      v0.27.14 | Pi plan progress survives compaction, Codex threads across rollout files, WSL browser setting, Mod+E edit mode
      v0.27.13 | Open a review on a specific base (--base, --diff-type), symlink containment on /api/doc, CI flake fix, Amp decision relay

      What's New in v0.27.24

      A small fix release for two code review problems found while recording a walkthrough of v0.27.23.

      Image previews stay in the all-files view

      When a diff changed images, the Before/After previews in the all-files view disappeared after a short scroll, and the view jumped ahead to the code files. The list measured each image card as having no height, so it dropped the cards as soon as they left the top of the screen. The cards are now measured at their real size, so they stay in place and keep their height while you scroll up and down. Scrolling a large diff is as fast as before.

      (#1652)

      PR comment previews open on the commented line

      In the PR Comments panel, the small code preview under a comment showed the start of the changed block, not the line that was commented on. For a new file that meant lines 1 to 5, even when the comment was on line 28. This was most visible on Bitbucket, but long blocks on GitHub had the same problem. The preview now opens on the commented line and the three lines above it, or on exactly the lines of a range comment. "Show full context" still shows the whole block.

      (#1651)

      Install / Update

      macOS / Linux:

      curl -fsSL https://plannotator.ai/install.sh | bash
      

      Windows:

      irm https://plannotator.ai/install.ps1 | iex
      

      Claude Code Plugin: Run /plugin in Claude Code, find plannotator , and click "Update now".

      Pi: Update @plannotator/pi-extension to 0.27.24 and restart Pi.

      OpenCode: Clear cache and restart:

      rm -rf ~/.bun/install/cache/@plannotator
      

      What's Changed

      • fix(review): PR comment previews open on the commented lines by @backnotprop in #1651
      • fix(review): keep all-files image previews in the virtual list by @backnotprop in #1652

      Full Changelog : v0.27.23...v0.27.24

    12. 🔗 HexRaysSA/plugin-repository commits sync repo: +2 releases, -1 release rss
      sync repo: +2 releases, -1 release
      
      ## New releases
      - [ida-mcp](https://github.com/hexrayssa/ida-mcp): 20260930.0.1
      - [ida-nexus](https://github.com/hexrayssa/ida-nexus): 0.13.1
      
      ## Changes
      - [ida-codemode](https://github.com/hexrayssa/ida-codemode):
        - removed version(s): 0.6.0
      
    13. 🔗 Rust Blog Announcing Rust 1.99.0 rss

      The Rust team is happy to announce a new version of Rust, 1.99.0. Rust is a programming language empowering everyone to build reliable and efficient software.

      If you have a previous version of Rust installed via rustup, you can get 1.99.0 with:

      $ rustup update stable
      

      If you don't have it already, you can get rustup from the appropriate page on our website, and check out the detailed release notes for 1.99.0.

      If you'd like to help us out by testing future releases, you might consider updating locally to use the beta channel (rustup default beta) or the nightly channel (rustup default nightly). Please report any bugs you might come across!

      What's in 1.99.0 stable

      extern "C" variadics

      Rust 1.99.0 stabilizes defining C-ABI variadic functions with "C" and "C-unwind" ABIs. Variadic functions defined this way use a variable argument list (...) and accept an arbitrary number of arguments. Rust could already call externally-defined variadic functions (e.g., libc::printf). With Rust 1.99, these functions can now be written in Rust itself:

      /// SAFETY: must be called with (at least) 2 i32 arguments.
      unsafe extern "C" fn sum(mut args: ...) -> i32 {
          // SAFETY: guaranteed by the caller.
          let a = unsafe { args.next_arg::<i32>() };
          let b = unsafe { args.next_arg::<i32>() };
          a + b
      }
      
      fn foo() -> i32 {
          unsafe { sum(0i32, 2i32) }
      }
      

      The type of ... is VaList, which is ABI-compatible with the C va_list type across targets. What types can be read from a VaList is guarded by the VaArgSafe trait.

      For more details on c-variadic functions, see the Reference. This release also stabilizes support for defining naked variadic functions with non-"C" ABIs, which must be written via inline assembly.

      Layout information from raw pointers

      This release settles the safety requirements for retrieving the size and alignment on raw pointers to both Sized (trivially safe, already possible on stable) and non-Sized types.

      This is done by stabilizing three functions:

      Recommend against round-trip unleaking after Box::leak

      While there are no changes to the language semantics in Rust 1.99, we have updated the documentation on Box::leak to recommend against patterns that later deallocate that memory. This was done because such code was found to have problematic interactions with current and future potential compiler optimizations, and is especially problematic with the upcoming stabilization of custom allocators. Instead, Box::into_non_null or Box::into_raw should be preferred.

      This guidance also applies to other leak functions in the standard library.

      Stabilized APIs

      Other changes

      Check out everything that changed in Rust, Cargo, and Clippy.

      Contributors to 1.99.0

      Many people came together to create Rust 1.99.0. We couldn't have done it without all of you. Thanks!

    14. 🔗 Console.dev newsletter tinyjs rss

      Description: Native webview desktop apps.

      What we like: Not Electron, so no bundled browser - uses the system native webview (macOS WebKit, Windows WebView2, Linux WebKitGTK). Native chrome with system API access e.g. files, sockets, processes. Small app bundles (runtime is 6MB). Everything is HTML inside.

      What we dislike: Everything is HTML, so whilst the wrapper feels native, the app isn’t.

    15. 🔗 Console.dev newsletter celld rss

      Description: Self-hosted durable objects.

      What we like: Each object (server) gets its own isolated SQLite database. Supports function execution, k/v, queues, cron triggers. Backed by object storage. Compatible with Cloudflare primitives and config. Can be distributed, with state managed in storage.

      What we dislike: Still early, with limited fleet management functionality e.g. auto-scaling, etc.

    16. 🔗 Filip Filmar rules_openxc7: Open-Source 7-Series FPGA Toolchains in Bazel rss

      rules_openxc7 brings open-source FPGA synthesis and place-and-route for AMD 7-series chips into Bazel. It drives Yosys, nextpnr-xilinx, Project X-Ray, and openFPGALoader through the same target interface as rules_vivado. This post covers the toolchain architecture, how the rules translate Xilinx constraints, reproducible bitstreams, and how to swap between Vivado and open tooling in one line.

      Why open tooling in Bazel

      Building for FPGAs usually means one of two extremes. You either install a 100-gigabyte proprietary vendor suite, or you maintain fragile shell scripts around open-source tools. rules_openxc7 avoids both problems.

    17. 🔗 New Music Releases Nightwish - Tribal (live in Amsterdam 2022) rss

      Nightwish - a new release is available:

      • 2026-10-01: Tribal (live in Amsterdam 2022) (Single)

      Amazon: Canada | Deutschland | France | United Kingdom | United States

      Visit muspy for more information.