🏡


  1. August 19, 2026
    1. 🔗 HexRaysSA/plugin-repository commits sync plugin-repository.json rss
      sync plugin-repository.json
      
      No plugin changes detected
      
    2. 🔗 anthropics/claude-code v2.1.236 release

      What's changed

      • Added ANTHROPIC_DEFAULT_MODEL environment variable: sets the model new sessions start on, while a /model pick still overrides it and persists across restarts (unlike ANTHROPIC_MODEL)
      • Added notify_when_idle to cross-session SendMessage: ask another Claude Code session on this machine to send one notice when it next goes idle — opt-in, one-shot, no polling (macOS and Linux)
      • Sandbox: on macOS, wildcard read-deny rules (e.g. **/.env) now take precedence inside allowed read regions, cover matched directories' contents, and can't be bypassed by renaming the denied file
      • Fixed clipboard copy, background housekeeping, background sessions, and local MCP logs breaking after the directory a session had switched into was removed (since 2.1.229)
      • Fixed the fullscreen renderer failing permanently after a single failed start: it now falls back to the classic renderer instead of exiting on every subsequent launch
      • Fixed the /model picker rendering taller than the terminal: it now shows only as many models as fit the window, with the rest reachable by scrolling
      • Fixed SendMessage calls being rejected when a malformed closing tag left the message text inside the summary field
      • Fixed unhandled promise rejections when a subprocess fails to start, for example powershell.exe on WSL with Windows interop disabled (regression in 2.1.234)
      • Fixed fullscreen mode sometimes not showing a newly sent message until the next update after the terminal was resized
      • Fixed a blank band that could remain above the prompt after clearing a multi-line prompt, and panes not repainting after resizing the terminal away and back, in fullscreen mode
      • Fixed the managed-settings approval prompt sometimes not appearing at startup while still capturing the first keypress as approval
      • Fixed terminal tab titles jumping in tmux (iTerm tmux integration): the title is now written only when its text changes instead of animating every 960ms
      • Fixed an unclear error when the cloud environments list came back empty or malformed
      • Fixed the Fable 5 first-time usage-credits prompt auto-selecting the fallback model after 60 seconds with no answer when using Remote Control
      • Fixed spinner tips never appearing, with a repeated background error, when the cached guest-pass reward in ~/.claude.json was malformed
      • Fixed skills hot-reload in SDK/VS Code sessions raising an error on every skills change after the session's working directory was deleted (2.1.229+)
      • Fixed self-hosted runner sessions released on idle, retire, or startup timeout occasionally resuming on another runner before the post-session hook had finished
      • Fixed the Clawd mascot's eyes and feet rendering unevenly in iTerm2 at some font sizes
      • Fixed occasional runaway session recaps: recap text (automatic and /recap) is now capped at 400 characters, cut at a word boundary
      • Improved startup performance: the session counter is now written in the background
      • Improved auto mode: Monitor allow rules are now set aside while auto mode is active, so Monitor commands are reviewed the same way Bash commands are
      • Improved auto mode on Bedrock, Vertex AI, and Foundry, and when telemetry is disabled: the classifier now uses the same defaults as on the Claude API, including severity-scored classification
      • Improved auto mode: the git status check can no longer be fooled by a repo's status.showUntrackedFiles=no setting into reporting a clean tree
      • Changed the /model picker to highlight only the newest model's name, so the highlight marks the new release rather than an arbitrary subset of the list
      • /goal: an idle session whose goal is parked behind long-running background work now checks in automatically after 30 minutes (then 1h, 2h) instead of waiting for you to return
      • /usage now shows the usage-credits spend row for Team and Enterprise members, and shows a capped row at 0% before anything is spent
      • SIGTERM in print/SDK mode no longer records an interrupted turn or synthetic tool denials before exiting; running commands are still terminated and the process still exits with code 143
      • Pressing Enter on a slash-command typo or a command unavailable in this session now reports it instead of running the closest fuzzy match; prefixes and aliases still run
      • Remote Control now marks a session offline within seconds when the CLI exits or its terminal closes
      • SendMessage now refuses further messages to a session up front once a rapid burst would exceed what that session's inbox accepts, instead of reporting them sent while they were dropped
      • Aligned the session title chip on the prompt border with the footer's right edge
      • Right-aligned footer items (goal indicator, session state, background agent status) and truncated notices now share a consistent right margin with the rest of the prompt area
      • [VSCode] Added screen reader support for the transcript: live announcements for replies, permission requests, errors, and status changes, plus per-turn heading navigation
    3. 🔗 r/LocalLLaMA Introducing Qwen3.8-27B Dynamic v3 Unsloth GGUFs rss

      Introducing Qwen3.8-27B Dynamic v3 Unsloth GGUFs | Hey everyone! We’re releasing new Qwen3.8-27B GGUFs with 10% higher accuracy for the same size. This uses a new version of Dynamic v3.0 Unsloth Dynamic V3 outperforms others by >10% on Div-300, KLD & more benchmarks. We also release 1-bit quants that retain 77% accuracy. Run on 8GB RAM. Some of you already saw we updated our quants a few hours ago. No, nothing was broken, nothing needed fixes (I don't know why people even said this since it's a complete fabricated story). This was purely an update to make them EVEN BETTER. We do not train on the imatrix calibration dataset, and we do NOT use QAT or QAD. Everything is done through post-training quantization. Our imatrix file used is available for the community to test, evaluate, and use. We encourage researchers and developers to create variations and fine-tunes of Qwen3.8 using our Unsloth quants/imatrix. You can read our over fitting analysis as well. Blog with all details and more benchmarks: https://unsloth.ai/docs/basics/dynamic-3.0-ggufs GGUF: https://huggingface.co/unsloth/Qwen3.8-27B-GGUF Enjoy! We also will be doing a new Unsloth Desktop update today: https://github.com/unslothai/unsloth We had A LOT of updates and will be introducing auto compaction, allowing external APIs to do tool calling and more. submitted by /u/danielhanchen
      [link] [comments]
      ---|---

    4. 🔗 The Pragmatic Engineer The Pulse: Grok’s CLI caught uploading all your local files to the cloud rss

      The Pulse: Grok’s CLI caught uploading all your local files to the
cloud

      Hi, this is Gergely with a bonus, free issue of the Pragmatic Engineer Newsletter. In every issue, I cover Big Tech and startups through the lens of senior engineers and engineering leaders. Today, we cover one out of four topics from a previous The Pulse issue. Full subscribers received the article below four weeks ago. If you 've been forwarded this email, you can subscribe here .

      Last week, xAI (Elon Musk's AI company, now part of SpaceX) released the Grok 4.5 model, built by Cursor (an acquisition), trained on SpaceX GPUs, and branded "Grok." It's a pretty good model; benchmarking as close in coding capability to Opus 4.8 and GPT 5.5 - while being 60-70% lower cost.

      Grok 4.5 can be used via API, but is easiest used via the Grok Build coding CLI. So, that's what many devs did. Some of them have noticed something really weird: the CLI is uploading all their local files in their working directories! Here's an independent AI safety researcher known as 'Cerblab' documenting what is happening (emphasis mine):

      "xAI's official Grok Build coding CLI (grok), on a normal consumer login, does three things worth documenting precisely:It transmits the contents of files it reads -- including a .env secrets file -- to xAI, verbatim and unredacted. The secret appears in two channels: the live model turn (POST /v1/responses) and a session_state archive uploaded and accepted (HTTP 200) via POST /v1/storage -- the endpoint the binary routes to the grok-code- session-traces GCS bucket (see section 5).It uploads the whole repository -- every tracked file's content plus git history -- independent of what the agent reads. Grok packages the workspace and uploads it via POST /v1/storage. Proven directly: on a real codebase, with the prompt "reply OK, do not read any files", Grok uploaded the entire repo as a git bundle (POST /v1/storage -> 200); git cloning the captured bundle recovers a file the agent was told not to open -- src/_probe/never_read_canary.txt -- with its unique marker verbatim, plus the full git history (appendix uploaded_repo.bundle). And it scales: on a 12 GB repo of never-read random files, /v1/storage moved 5.10 GiB, all HTTP 200 (truncated mid-stream), while the model-turn channel moved just 192 KB -- a ~27,800× ratio that pins the upload to the codebase, not to what was read. No storage upload failed; the only non-200s were a model-usage quota (402/429) on /v1/responses and one unrelated 404 -- not a storage size cap.The storage destination is a Google Cloud Storage bucket, grok-code-session-traces (not AWS S3) -- named verbatim in the binary and in a captured metadata.json (gs://grok- code-session-traces/…). I did not find this mechanism surfaced in the CLI's install/quickstart materials (not an exhaustive docs audit -- §7), it is active by default, and disabling "Improve the model" does not turn it off (/v1/settings still returned trace_upload_enabled: true; §6).

      None of this proves xAI trains on the data -- that is a policy question addressed in [section] 6. What is proven is transmission, acceptance, and storage."

      There's so much wrong with this approach! To name a few:

      • Sending your codebase over the context window is not normal. All AI agents send over their context window to the server that runs the LLM. That's where tokenization happens and the context is appended to the session. This means that other AI agents send over some part of the code that they read in their context window.
      • No need to send over the source code to index it. Indexing the codebase is important for efficient code lookup, and we covered how Cursor does this in a privacy-conscious way by indexing a user's codebase locally, creating embeddings, and sending those embeddings to the server. Cursor's server does not store any of the user's codebase though. Except that Grok CLI transferred all users' codebases to a cloud bucket! Given Cursor and Grok are now combined as part of SpaceX, it 's a real head scratcher why the Grok CLI isn't doing what Cursor always has.
      • Sending over unencrypted .env files is reckless. Local .env files store secrets, database access tokens, service access tokens, and more. These are sensitive pieces of information that need to be handled with care. If transmitted, they should be encrypted at the very least. Grok / SpaceX storing them on the GCP storage bucket - likely unencrypted - is flat-out unacceptable and reason enough for any sensible company to ban usage of Grok CLI.
      • No good reason to upload git history. Sure, seeing Git history could be helpful when training an AI model.
      • It 's malicious to not tell devs anything about it. Developers using Grok CLI haven't been asked to opt into this data upload, nor notified of it. Most evidently had no idea this has been happening, and are understandably furious after Cerblab's writeup went viral.

      SpaceX throws devs "under the bus"

      Caught red-handed, Grok CLI disabled file uploads with a remote feature flag. AWS engineer Wes Eklund started tracking upload functionality with the CLI, finding that the uploads suddenly stopped due to a feature flag being flipped by the Grok team, pausing data collection.

      But the code functionality to stream all local files to the server, unencrypted, remained present in the CLI, even in later updates. Read an in-depth analysis by Wes.

      SpaceX 's official response was pretty laughable, not explaining why .env files and .git history were uploaded, and adding that enterprise customers with zero data retention (ZDR) enabled were the only ones unaffected by underhanded, secret uploading of users' local files:

      The Pulse: Grok’s CLI caught uploading all your local files to the
cloudTranslation: "Enterprise users with ZDR turned on were not impacted. To everyone else: we didn't tell you about this hidden/privacy command, but it's your fault. Source: SpaceX

      My initial reaction is what a condescending response by SpaceX, swiftly followed by the question: does Grok CLI even have enterprise customers? It might be few, given how reckless the team evidently is by uploading unencrypted secrets! SpaceX CEO Elon Musk chimed in with a post that seemed to be almost trolling angry users. He wrote:

      "SpaceX policy regarding data retention.

      It is actually helpful for debugging issues if we can retain some amount of data, so allowing this would be appreciated, but your privacy settings are always respected."

      This makes it worse because SpaceX has been secretly uploading far more data than is "useful for debugging!" Uploading the git history and sensitive .env files is not , in any way, useful for debugging. Also, Grok/SpaceX did not upload "some amount of data"; it uploaded every last file it could find in your local folder.

      The developer community is justifiably upset to read SpaceX and Musk pretending that Grok has only uploaded scraps of data purely for debugging purposes. This time, even fans of SpaceX and Grok are speaking out against Musk and his company's behavior. AWS engineer Wes Eklund:

      "Elon, firstly, huge fan of everything you work on. Completely understand the need for some trace data to improve customer experiences.

      From what I've researched, it seems to be much more than just trace debugging issues. It seems to be entire code repos with sensitive information just collected entirely.

      Your Google Cloud blob storage must have petabytes of code repos from us.

      Not ideal."

      Sam Altman pushes Grok to open source Grok CLI

      OpenAI CEO, Sam Altman, also posted, using a term Musk often employs when commenting on things he disapproves of in society, and hinting at the benefits of Codex which doesn't upload your whole local filesystem, or mess with .env files and the git history. The Codex harness is open source, so secretive file-upload functionality would be visible in the source code:

      The Pulse: Grok’s CLI caught uploading all your local files to the
cloudAltman uses Musk 's trademark "concerning" remark against him. Source: Sam Altman

      Musk clearly read it and responded a few hours later:

      The Pulse: Grok’s CLI caught uploading all your local files to the
cloudMusk committing to open sourcing Grok CLI. Source: Elon Musk

      A day later, (15 July), SpaceX did indeed open source Grok CLI. Altman's comment seemingly hit home. Also, SpaceX has stated it is deleting data from its servers, writing:

      "We disabled default retention for all Grok Build users starting on July 12th. Additionally, we are deleting all coding data that was previously retained, ensuring every user's preferences are respected. With these steps, Grok Build goes beyond other major coding products to protect user privacy."

      The open sourced repo is rushed, unsurprisingly. Kernel engineer Elliot Arledge used the repo and reported what he found:

      "Out of the box, cargo test --workspace doesn't compile. 190+ errors, all one bug class: cross-crate test helpers hidden behind #[cfg(test)], which Bazel's per-target test builds tolerate but Cargo doesn't, because a dependency is never compiled in test mode. The default-bazel feature is declared in ~20 manifests and wired to nothing. Ungating the benign helpers and putting the signing-key test seam behind a dev-dependency-only feature gets the suite running for the first time on the published tree: 24,663 passed, 28 failed, and every one of the 28 is a pre-existing bug the broken build had been hiding (tests reading your real ~/.claude/settings.json, a tool missing from the registry, macOS /var symlink breakage, one Theme::current() race).

      Questions for the team:

      Did anyone try downloading this repo and running it before publishing?"

      In fairness, it's clear from the outside that the dev team was instructed to open source the repo ASAP, and did just that. I would expect improvements to allow the developer community to compile the repo and run tests will come later.

      On a side note, I wonder what the mood within Grok is. A mandate to open source the product - that was not built to be open sourced! - and doing it in a couple of hours hints what it's like to work there. Some folks doubtless would find it thrilling and a big challenge, while others likely find it pretty stressful having the "Eye of Sauron" (the CEO) upon them!

      Can companies trust Grok now?

      To hand it to Grok/SpaceX, the last 24 hours of this incident saw some impressive execution. In contrast, everything that took place beforehand screams "amateur hour":

      • Did no one in teams that wrote the functionality to upload the complete local codebase raise concerns that this would be unacceptable to devs?
      • How did uploading secrets without encryption not raise alarm bells?
      • Why greenlight any of it without opt-in and do it in a secretive manner?

      For sensible companies generating revenue with software, vendors who are allowed to access their codebase are limited to those that can be trusted. Grok has demonstrated it is unprepared to handle codebases with security fundamentals in mind.

      There's now frantic backpedaling and rapid open sourcing, but my impression is that this is simply because Grok/SpaceX was caught red-handed, and understands the risk of losing enterprise contracts caused by this flagrant breach of trust. This is why SpaceX's communication mentioned that enterprise customers with zero data retention have not been impacted!

      Trust is earned in drops and lost in buckets; SpaceX/Grok will likely learn this. By secretly uploading codebases and secrets, Grok CLI revealed itself as an untrustworthy coding agent that's a risk to use. Meanwhile, all the major competitors - Claude Code, Codex, OpenCode, Gemini CLI - have never violated user trust this way.

      Grok CLI can rebuild trust, but it'll likely take years of no security-related incidents to prove they are an open source-first product (that they were not, just a day ago!), and also demonstrate that they care about "normal" developers, and not just enterprise clients with ZDR turned on.

      This incident is a good reminder of why planning and process can slow shipping speed, but increase revenue generation. I would wager that the Grok team has shrunk SpaceX's enterprise subscription prospects for the foreseeable, in the name of saving a few hours on security reviews, learning what other coding harnesses do with codebase uploads, or even just asking the Cursor team!

      I predict Grok/SpaceX will have to offer very high usage limits inside the Grok CLI to convince devs to take a risk on running this software on their system. And they will have to undercut OpenAI and Anthropic API pricing massively for any security team to greenlight use of a CLI that just last week was sending .env secrets unencrypted to their GCP buckets.

      Of course, SpaceX/Grok will be just fine as it has the capital to fix things. It will now just be a lot more expensive and time-consuming to fix something that was likely caused by a few engineers wanting to make debugging easier!

      Read the _full _The Pulse issue__ , or check out this week 's The Pulse . The full issue additionally covers:

      1. New trend: concern about massive increase in code review load. Top of mind for engineering leaders: what to do about the ever-growing code review load, and how devs are starting to review code less thoroughly than before? Many questions, but few proven solutions. Send comments about what you see working.
      2. Are more devs at enterprises upset about enterprise pricing by AI labs - and does it matter? I got a message from a reader baffled to learn their company pays 20-30x the price for tokens than their own $20/month Claude Code / Codex subscription. It may show how valuable AI coding tools are.
      3. Linux creator: AI "clearly useful." Inside the Linux kernel maintainers group, the discussion veered onto whether Linux should consider banning AI contributions, similar to how some FOSS projects have done so. Linus Torvalds weighed in and made it clear that AI is useful, everyone should decide whether to use it, but no one is allowed to tell others what tools to use. Given AI is an increasingly capable tool, it would be foolish to not use it as such.

      Read the full The Pulse

    5. 🔗 @binaryninja@infosec.exchange Sidekick can use the terminal now! With the new run_terminal_command tool, mastodon

      Sidekick can use the terminal now! With the new run_terminal_command tool, agents can jump into bash or PowerShell right alongside their Binary Ninja analysis. Unpack an archive, run binwalk, poke through git history, curl a URL, or inspect files that never even made it into Binary Ninja. Commands still require your approval by default. https://sidekick.binary.ninja/blog/sidekick-26-1-a-proper-home-for- sidekick/#sidekick-can-use-the- terminal

    6. 🔗 MetaBrainz Growing pains: An update on the ListenBrainz service status rss

      TL;DR: The growth of ListenBrainz has caught up with us, and our limited team is working on replacing central parts of our infrastructure. In the meantime, many features are unstable.

      We are victims of our success. While in the long term this is a good problem to have, the sharp increase in users over the past year and a half has left our infrastructure cracking at the seams.

      The good news: Your listens are being stored and imported without issue, even if there are delays. The core service is not compromised and we ask that you please keep submitting your listens. Your stats and playlists will be back!

      However, we know that every other week your statistics, weekly playlists, and other features fail to generate for everyone, and cause crashes. We know it is frustrating, and we share that feeling.

      While we are aware of the issues, fixing them is far from simple and requires us to completely rework all the crucial parts of our infrastructure.

      In the interest of transparency, here are our main issues and what we are doing to fix them:

      How many listens?

      Our database dumps have become too big for us to process.
      We went from 0 to 1 billion listens in 7 years - then to 2.5 billion in the next one and a half years.

      Generating and copying our (now huge) database dumps causes crashes as our limited servers run out of memory and disk space.

      We are working on improvements, but for each change we need to wait two days to be able to run tests.

      In addition, we are moving away from TimescaleDB, a Postgres database extension for time-series data, in favour of vanilla Postgres with a new partition scheme.

      We found that TimescaleDB was not adapted for our use case of working with historical imports or deleting listens and users (and all their listens), as well as large gaps between listens, all causing some very slow queries.

      Where are my stats, goddammit?

      Our statistics and playlist calculation infrastructure, a Spark cluster of 5 servers, is running out of memory and crashing, from one task or another. This used to happen once every few months but is now a weekly occurrence.

      This is the issue which is breaking stats, playlists, user similarity, unlinked listens, fresh releases and more.

      We are moving to using Clickhouse instead for all statistics calculations, which will free up the Spark cluster to be used for generating playlists and other tasks.
      This move is taking some time, as a single person in the team carries the responsibility of rewriting essential code and testing everything carefully.

      This will eventually open the door to requesting stats for an arbitrary time range instead of being limited to this and last week/month/year, a hotly requested feature that is not possible with our current system.

      The scraping situation

      To make matters worse, the entire internet is being bombarded by unscrupulous bad actors (looking at you, AI companies) that don't follow the rules and try very, very hard to evade any measure meant to limit them.

      They rent botnets of millions of residential IPs so they can scrape our APIs and websites incessantly, over and over again, while evading detection, all for data that they could download for free.

      They cause surges of 5x the usual traffic across all our projects and slow everything down on our resource-constrained infrastructure.
      It also forces us to spend time dealing with these DDOS- like surges instead of working on our other pressing issues.

      For ListenBrainz specifically, we have had to disable some features/endpoints that during scraping waves made the website completely unreachable for everybody.

      But wait, there's more…

      The loss of our founder in late February was big blow to our team.

      Rob was one of the custodians of Listenbrainz infrastructure, but also a central ListenBrainz team member.

      We have had to cross-train our ListenBrainz dev team -it is only three of us- to deal with infrastructure and other new aspects, as well as reorganize priorities to deal with the day-to-day operations while the foundation was in the process of hiring a new executive director.

      We have been so greatful for the wonderful patience and kindness shown to us by you, our users and community, as we work through these growing pains. Keep submitting listens and let's grow together!

      • your ListenBrainz Team
    7. 🔗 syncthing/syncthing v2.1.4-rc.1 release

      Major changes in 2.1

      • Devices and folders can now be grouped in the GUI by setting the new
        group attribute.

      • HTTP and HTTPS proxies with support for CONNECT can now be used, in
        addition to the existing support for SOCKS proxies (the environment
        variable all_proxy=https://...).

      • Block indexing can be turned off for folders where it's more desirable to
        optimise for reduced database size and overhead than minimal transfer
        size (the blockIndexing attribute on folder configuration).

      • GUI login session duration can be configured to be longer or shorter than
        the default one week, or set to infinitely long. The cookie path can also
        be adjusted. (The sessionCookieDurationS and sessionCookiePath
        attributes in the GUI configuration.)

      This release is also available as:

      • APT repository: https://apt.syncthing.net/

      • Docker image: docker.io/syncthing/syncthing:2.1.4-rc.1 or ghcr.io/syncthing/syncthing:2.1.4-rc.1
        ({docker,ghcr}.io/syncthing/syncthing:2 to follow just the major version)

      What's Changed

      Fixes

      • fix(model): correctly handle receive-only changed directories (fixes #8004) by @calmh in #10843
      • fix(api): correctly return metrics, support bundle (fixes #10847) by @calmh in #10849
      • fix: open GUI when relaunched instead of printing error (fixes #10727) by @calmh in #10852

      Other

      Full Changelog : v2.1.3...v2.1.4-rc.1

    8. 🔗 WerWolv/ImHex Nightly Builds release

      Nightly

      97b91d6 Changelog

      • patterns: Update pattern language
      • build: Update libwolv
      • build: Upgrade GCC
      • build: Make sure Snap doesn't try to use host libc
      • fix: Remove unused lang entry
      • build: Update libwolv
      • build: Update libwolv
      • feat: Add support for 32 bit and 64 bit time_t on all platforms
      • fix: Settings getting reset after using cli
      • build: Fix build issues
      • patterns: Update pattern language
      • feat: Rework pattern language #pragma matcher system, add filename and provider type matcher
    9. 🔗 PrimeIntellect-ai/prime-agent Beta (v0.7.3-beta.518.1.f8f0036) release

      Automated beta build from main (f8f0036cc2da1a640aad990ae8dcb7c4820ce32e).

    10. 🔗 Armin Ronacher What Is Reasoning rss

      A few weeks ago a paper was shared that showed how to extract reasoning traces from closed-weight models. Together with online discussions about tricking models into leaking them, it made me investigate it more out of curiosity. Twitter seems full of half-truths and confusion about how this works, so perhaps this helps some to understand what is happening.

      Hiding Traces

      Reasoning traces are usually hidden from us. We have lamented this, but mostly have to accept it. Open-weight models thankfully reveal them, and from their behavior you can see that their traces can be long and confusing. This is probably a good reason to separate them from what is normally shown to users.

      At minimum, UIs need to detect them. The industry has done a good job at making reasoning traces sound special and exotic, but they really are just text: the model is trained to emit its thinking into a scratchpad as part of its response, before its final answer.

      GPT-OSS's Harmony response format makes this easy to see:

      <|channel|>analysis<|message|>
      I need to work this out ...
      <|end|><|start|>assistant<|channel|>final<|message|>
      The answer is ...
      <|return|>
      

      The markers are special tokens, but the reasoning between them uses "the same text" as the final answer (just that GPT chain-of-thought text sounds really funny). When the model samples the analysis channel token, a parser routes the following text into a separate stream exposed through the Responses API. For closed models, presumably a simple model redacts and summarizes it.

      Reasoning Effort

      How much budget goes to reasoning? Earlier APIs exposed reasoning token budgets, making it seem like a property of the sampling process. In reality, reasoning effort is baked into the system prompt. GPT-OSS puts this into the system prompt:

      Reasoning: low
      

      That's it. Training produces the resulting behavior, such as emitting the token sequence that switches to the analysis channel. This also explains why changing the effort invalidates the KV cache. I think closed GPT models call reasoning effort "juice," since you can ask most models how much juice they have.

      In DwarfStar for DeepSeek with max reasoning this is added to the system prompt:

      Reasoning Effort: Absolute maximum with no shortcuts permitted.
      You MUST be very thorough in your thinking and comprehensively decompose the
      problem to resolve the root cause, rigorously stress-testing your logic against
      all potential paths, edge cases, and adversarial scenarios.
      

      Don't Think

      The destination of reasoning tokens is therefore a learned convention: the model is trained to keep scratch work out of the final channel. Trick it into thinking it is in that channel and it may leak tokens. We have even seen older models, when thinking is disabled, reason into the bash tool and echo their thoughts to /dev/null.

      So in some sense the only "special" behavior for some models is not to think. That at times is done by "mechanically" removing the model's usual ways to think. In DwarfStar, disabled thinking uses the prefill </think>, while enabled thinking uses <think>, which are the tokens that close and start thinking. GPT-OSS doesn't prefill but lets the model decide either way on its own.

      But presumably, some inference APIs prefill the opening token when reasoning is enabled, so the model never samples it itself and might prevent the sampling of the reasoning token when disabled since it can be trivially detected. This may explain why a custom think tool can trick models into putting some reasoning where it should not go — but only when native reasoning is disabled.

      Fun fact: this blog post triggered safey checks

      Hilariously enough I was unable to use GPT 5.6 terra for spell and grammar checking on this blog post because of safety filters. Had to switch to Kimi.

      GPT-5.6-terra refusing to spell-check this blog
post

    11. 🔗 Ampcode News MCP in Orbs rss

      You can now connect remote MCP servers on ampcode.com and use them in orbs, the TUI, and with Puck.

      The MCP settings screen showing preconfigured MCP servers

      To add an MCP server:

      1. Visit ampcode.com/settings/mcp-servers,
      2. Select a pre-configured server or click "Add MCP Server"
      3. Log into the server using OAuth,
      4. Start a thread in an orb, the TUI, a runner, or talk to Puck:

      Amp supports connecting hosted MCP servers using Streamable HTTP with OAuth or Bearer tokens, personal and workspace configuration of MCP servers, and the use of MCP-provided tools.

      MCP Apps, Resources, and Prompts are not supported.

    12. 🔗 Ampcode News Pass the Orb to the Left Hand Side rss

      You can now bring your team into your orb.

      Use @ to tag members of your team, and they'll be able to view, drive, and send chat messages to the thread.

      We've been having a lot of fun with this feature recently. Here are a few examples:

      Thorsten Ball passes a request for an orb sticker from Tim
      Tim Lucas mentions Lewis Metcalf to ask for feedback on a code diff
      A teammate mentions Tim to ask for a documentation review
      A message copies Camden Cheek into a thread
      A teammate mentions Rocko Reager to ask for help with a sidebar bug
      Tim Lucas mentions Rocko Reager to ask if he saw a message
      Teammates coordinate synchronized orb animations in a thread
      Tim Culverhouse mentions Lewis Metcalf before asking Amp to start work
      Thorsten Ball mentions Brett to ask about shipping an orb navigation change
  2. August 18, 2026
    1. 🔗 IDA Plugin Updates IDA Plugin Updates on 2026-08-18 rss

      IDA Plugin Updates on 2026-08-18

      New Releases:

      Activity:

      • claude-marketplace
      • diaphora
        • f997fd34: Update README.md
        • 31282dee: Merge branch 'master' of ssh://github.com/joxeankoret/diaphora
      • disrobe
        • ebdb7657: sleigh: build the unsupported effect name from a sanitised mnemonic s…
        • 94e0f342: binfmt: recover stuffit 5 entry dates, type, creator and finder flags
        • 7df8d145: dotnet: record how a nativeaot fixture project must be named before i…
        • 0cd5a4d2: cli: grade rar recovery through auto at one and four workers
        • d63e8083: binfmt: recognise stuffit 5 by its own signature and extract its memb…
        • f45554bf: binfmt: emit rar members as container chain children
        • bc019223: dotnet: record whether a nativeaot signature came from managed metada…
        • d4fcd174: binfmt: track the rar fixtures in the ignore list and correct the stu…
        • 557408e2: hermes: recover labelled loop exits and read the one-byte string key tag
        • afd8c25e: binfmt: recover rar 2.9/3.x filter records and mixed lz and ppmd blocks
        • eec74559: binfmt: decode stuffit method 5 forks and stuffit 5 containers on the…
        • da660ddd: arm32: choose the decode mode from the elf mapping symbols instead of…
        • 1f318d00: native: keep aarch64 d-register classification off callee-saved spill…
        • ef62de45: native: exercise the aarch64 simd memory forms on the decompile comma…
        • 20e738af: binfmt: decode lzh pm1 and pm2 members and count the aarch64 nan grad…
        • 4ff291f9: sleigh: vendor the aarch64 simd specification whole and lift the rema…
        • 6c9d4db9: native: recover aarch64 stack floating parameters and grade thirty sc…
        • 30e1bdbe: py: keep the guard on a with region a conditional encloses and republ…
        • dcf62039: recovery: lift aarch64 scalar floating transfers and reattach void an…
        • 445c68ce: native: restore four aarch64 recoveries and report a non-equivalent p…
      • distro
        • a20f55f4: restrict-egress 0.10.1: version bump for the diagnose allow-set fix
        • 29efc0cf: restrict-egress: read allow sets after resolving in diagnose
        • e90be796: restrict-egress 0.10.0: couple DNS answers to the allow set; add diag…
      • eject_idb
        • 0027b134: ci: bundle the standalone signaler binaries into a single tools zip
        • ee90498e: Maintenance release (v0.0.4)
      • ida-free-mcp
        • fb3b0c89: fix(decompile): return the requested function, not a stale pseudocode…
      • IDA-PRO-MCP
        • 0211e84b: chore: add repository URLs and author metadata
        • f6b1001c: Initial commit: IDA Pro MCP server with 87 tools
      • IDAPluginList
        • cc791482: chore: Auto update IDA plugins (Updated: 19, Cloned: 0, Failed: 0)
      • mytools
        • bb7ec0c6: Update rootfs-tools skill and macOS workflow
      • twdll
        • 4e80db7c: feat(attila): add ARTSET_SCRIPT_INTERFACE and character portrait API
        • ce1aac80: docs(cai): create Campaign AI Telemetry and Decision Intelligence Arc…
        • 3a7ebe40: fix(cai): use exact DWARF enum names for OCCUPATION_DECISION
        • 2ebaede2: feat(cai): add Campaign AI occupation telemetry and decision logging
        • 5915a726: feat: add region religion breakdown api and standardize list method n…
        • 5a7bade4: fix: preserve proportional health and sync bodyguard in ConvertUnit
    2. 🔗 anthropics/claude-code v2.1.235 release

      What's changed

      • Added an optional spellcheck setting that underlines misspelled words in the prompt input as you type, using your installed aspell, hunspell, or ispell
      • Fixed whole-prompt-cache invalidation when a language server disconnected or reconnected mid-session
      • Fixed nested markdown list items misaligning at depth 3+ and added a hanging indent to wrapped list items in the terminal UI
      • Fixed prompt input highlights (slash commands, keywords, mentions) appearing shifted by one or more characters in some multi-line prompts
      • Fixed Shift+Tab inside the permission prompt's comment field approving the edit and granting session-wide edit permission instead of closing the field
      • Fixed the Agent tool advertising a general-purpose default in sessions where that agent is unavailable: an omitted subagent_type there now gets a clear error listing the available agents
      • Fixed notebook cell delete/replace approval dialogs silently omitting the existing cell content when the notebook or cell could not be read; the dialog now says why
      • Fixed slash commands run while Claude is responding showing HTML entities instead of the actual characters
      • Fixed the prompt footer not showing the "Update installed" restart notice after a background auto-update
      • Fixed the expanded task list (ctrl+t) always starting collapsed when resuming or relaunching into a session that still has open tasks
      • Improved memory and CPU usage while cloud sessions such as /ultrareview or /autofix-pr run in the background — their event streams are no longer re-scanned and re-rendered on every update
      • Improved permission dialogs: display text and "don't ask again" options now always match what a grant would cover, and "don't ask again" is withheld when contents cannot be fully displayed
      • Improved the embedded grep in native macOS/Linux builds: pathological patterns now fail fast instead of exhausting memory, and -m N with -A/-C prints correct context
      • Improved the context-limit error to say when auto-compact is off and point to /config to re-enable it
      • Vim mode: NORMAL mode and cursor position are now preserved when toggling the detailed transcript (ctrl+o) or closing a panel
      • Dialogs: arrow keys and Enter pressed in quick succession now select the option you navigated to instead of the previously highlighted one
      • SendMessage now refuses messages too large for cross-session delivery up front instead of silently dropping them
      • Remote Control: claude rc now applies the same enterprise-gateway availability check as interactive startup
      • [VSCode] Fixed focus jumping between open Claude tabs on its own when a window with several Claude panels is restored or reloaded
    3. 🔗 uswds/uswds USWDS 3.14.0 release

      What's new in USWDS 3.14.0

      Features

      Package | A11y | Breaking | Markup change | Description
      ---|---|---|---|---
      usa-accordion, uswds-core | Yes | Yes | Yes | Added a left-aligned expand/collapse icon option. A new $theme-accordion-icon-position setting (default: "start") lets teams set icon placement globally. The .usa-accordion--icon-start and .usa-accordion--icon-end modifier classes support per-instance placement. This improves discoverability for users viewing content at high magnification or zoom levels. Thanks @HopeTurnerUSCIS, @rosamundtgov, and @jeana-adhoc! (#6789)

      ✏️ Teams should verify layout at common zoom levels for the new default accordion behavior.
      usa-breadcrumb | Yes | Yes | - | Breadcrumbs now wrap by default. The previous truncation behavior is now opt-in via the new usa-breadcrumb--truncate modifier class. Thanks @AKnassa! (#6722)

      ✏️ Teams should confirm breadcrumbs display as expected and add usa- breadcrumb--truncate if they want the old truncation behavior.
      usa-range | Yes | - | Yes | Added a visible hint to the range slider. A new usa-hint element with the text "Move the slider to change the value" is added above the slider so sighted users receive the same guidance that screen reader users already had. Thanks @ravitejapioneerblaze-code! (#6673, #6811)

      ✏️ Teams should pull in the updated markup.
      usa-range | Yes | - | - | Improved range slider border visibility. The border is now 2px and uses the base-darker theme token. A focus ring is also added to the slider input. Thanks @manichandra! (#6659)
      usa-date-picker | Yes | - | Yes | Addedaria-current="date" to today's date button. Assistive technologies can now programmatically identify the current date in the calendar widget. Thanks @daresTheDevil! (#6593)
      usa-file-input | - | - | - | Error border now uses theerror-dark token. The file input error state previously used secondary-dark, which could show the wrong color in projects with distinct secondary and error palettes. Thanks @manichandra! (#6669)

      Bug fixes

      Package | A11y | Breaking | Markup change | Description
      ---|---|---|---|---
      usa-modal | Yes | - | - | Closing a modal now always restores screen reader access to the page. If the element that opened the modal had left the document by the time the modal closed, page content kept aria-hidden="true" and stayed invisible to assistive technology until reload. Thanks @vssinghh! (#6786)
      usa-modal, uswds-core | Yes | - | Yes | Fixed modal content being read twice by screen readers. The default focus target has changed — on open, focus now moves to the first enabled button in the modal footer, or the first enabled button in the modal if no footer button is present. FocusTrap no longer uses autoFocus. (#6703)

      ✏️ Teams should verify modal focus lands where expected after this update.
      usa-memorable-date | Yes | - | Yes | Added per-field hints to the memorable date component. The component previously used a single shared hint referenced by all three fields via aria-describedby, causing screen readers to repeat the full instruction block on every field focus. Each field now has its own targeted hint. The visible group hint remains for sighted users with aria-hidden="true". (#6725)

      ✏️ Teams should update to the new per-field hint markup. Teams supporting other languages should update their hint strings.
      usa-file-input | Yes | Yes | Yes | Removed the drag instruction on mobile and coarse-pointer devices. The file input previously showed "Drag file here or choose from folder" on all devices, including mobile where drag-and-drop isn't a practical interaction. Coarse-pointer devices now show and announce "Choose from folder" only. Thanks @manichandra! (#6660)

      ✏️ Teams should check for layout changes and update any accessibility tests that assert the old exact instruction text.
      usa-character-count | Yes | - | - | Deferredaria-live to prevent iOS VoiceOver from announcing character count on page load. Live updates continue to fire normally during typing. Thanks @daresTheDevil! (#6595)
      usa-character-count | - | - | - | Fixed the label selector so it correctly finds the associated label. Thanks @ealexhaywood! (#6385)
      usa-footer | - | Yes | - | Restricteddata-tag to heading elements. The footer's big link list was vulnerable to XSS through unsanitized data-tag values. Non-heading elements now gracefully fall back to h4. Thanks @IHIutch! (#6674)

      ✏️ Teams should review any footer data-tag values that aren't heading elements.
      usa-banner, uswds-core | - | - | - | Fixed toggle so it resolvesaria-controls from the component's root node. This allows banner toggles to work correctly when used inside a shadow root. Thanks @arpitjain099! (#6714)
      uswds-core | - | - | - | Fixed the language selector Escape key handler. Thanks @arpitjain099! (#6713)
      uswds-core | - | - | - | Guarded the keymap against non-keyboard events from datalist selections. This prevents a console error when a user selects an option from a datalist. Thanks @vijaygovindaraja! (#6594)
      usa-table | - | - | - | Restored the row header border in borderless tables. Row-scoped body header cells now keep their top border so row headers don't appear visually disconnected. Thanks @manichandra! (#6661)
      usa-table | - | - | - | Fixed.usa-sr-only table caption causing heading-row border collapse. Thanks @IHIutch! (#6633)
      usa-time-picker | - | - | - | Added missing combobox style dependency to the time picker package. Thanks @IHIutch! (#6634)
      usa-input, usa-textarea, usa-range, usa-combo-box, usa-input-prefix-suffix, usa-select | - | - | - | Setbox-sizing: border-box on the %block-input-styles mixin to prevent overflow. Components in host environments that reset global box sizing no longer overflow their containers. Thanks @VenkateshAddala! (#6736)

      ✏️ Teams should verify these elements display correctly in projects with custom global box-sizing resets (e.g. those who set $theme-global-border- box-sizing: false).
      usa-in-page-navigation | - | - | - | Standardized the component's enhancement guard to usedata-enhanced in line with other USWDS components. (#6688)
      uswds-core | - | - | - | Fixed ink assignment referencing the wrong variable. Custom ink colors in a project's theme were referencing the wrong system token. Thanks @nektro! (#6651)

      ✏️ Teams should verify that custom ink colors in their theme render as expected.

      Guidance changes

      Alert

      The alert component page now recommends that alert headings start with the alert type to improve clarity and urgency of the message, both for accessibility and general usability, as well as adding extended guidance to help teams use the correct alert type. Thanks @jeana-adhoc and @rosamundtgov! (#3288)

      Markup changes

      Memorable date

      The memorable date component's three fields now each have their own aria- describedby hint instead of sharing a single
      group hint. Teams who've copied the memorable date markup should update to the per-field pattern:

       <fieldset class="usa-fieldset">
         <legend class="usa-legend">Date of Birth&lt;/legend&gt;
      -  <span class="usa-hint" id="mdHint">For example: January 19 2000&lt;/span&gt;
      +  <span class="usa-hint" aria-hidden="true" id="memorable-date-hint">
      +    Select a month. Enter 1 or 2 digits for the day and 4 digits for the year.
      +  &lt;/span&gt;
         <div class="usa-memorable-date">
           <div class="usa-form-group usa-form-group--month usa-form-group--select">
             <label class="usa-label" for="date_of_birth_month">Month&lt;/label&gt;
      -      <select class="usa-select" id="date_of_birth_month" name="date_of_birth_month" aria-describedby="mdHint">
      +      <span class="usa-hint usa-sr-only" id="memorable-date-month-hint">Select a month from the dropdown.&lt;/span&gt;
      +      <select class="usa-select" id="memorable-date-month" name="memorable-date-month" aria-describedby="memorable-date-month-hint">
               ...
             &lt;/select&gt;
           &lt;/div&gt;
           <div class="usa-form-group usa-form-group--day">
             <label class="usa-label" for="date_of_birth_day">Day&lt;/label&gt;
      -      <input class="usa-input" aria-describedby="mdHint" id="date_of_birth_day" name="date_of_birth_day" ... />
      +      <span class="usa-hint usa-sr-only" id="memorable-date-day-hint">Enter 1 or 2 digits for the day.&lt;/span&gt;
      +      <input class="usa-input" aria-describedby="memorable-date-day-hint" id="memorable-date-day" name="memorable-date-day" ... />
           &lt;/div&gt;
           <div class="usa-form-group usa-form-group--year">
             <label class="usa-label" for="date_of_birth_year">Year&lt;/label&gt;
      -      <input class="usa-input" aria-describedby="mdHint" id="date_of_birth_year" name="date_of_birth_year" ... />
      +      <span class="usa-hint usa-sr-only" id="memorable-date-year-hint">Enter 4 digits for the year.&lt;/span&gt;
      +      <input class="usa-input" aria-describedby="memorable-date-year-hint" id="memorable-date-year" name="memorable-date-year" ... />
           &lt;/div&gt;
         &lt;/div&gt;
       &lt;/fieldset&gt;
      

      File input

      The file input no longer renders the drag instruction on coarse-pointer or mobile devices. On those devices, the
      instruction now reads "Choose from folder" instead of "Drag file here or choose from folder". For fine-pointer devices,
      the text remains "Drag file here or choose from folder." Teams with tests that assert the exact instruction text should
      update those tests.

      - Drag file here or choose from folder
      + Choose from folder
      

      Accordion icon alignment

      The default alignment for the accordion toggle icon switches to the left for better accessibility for users who zoom or
      use screen magnification. Teams can now use the usa-accordion--icon-start or usa-accordion--icon-end modifier to
      left-align or right-align the expand/collapse icon at the instance-level respectively. Icon position can be set globally
      with the $theme-accordion-icon-position Sass setting. The default behavior ("start" / left-aligned) is changed from
      v3.13.0, and the historical behavior can be preserved with $theme-accordion- icon-position: "end".

      - <div class="usa-accordion">
      + <div class="usa-accordion usa-accordion--icon-start">
      

      or

      - <div class="usa-accordion">
      + <div class="usa-accordion usa-accordion--icon-end">
      

      Or set globally in your theme:

      + $theme-accordion-icon-position: "start";
      

      or

      + $theme-accordion-icon-position: "end";
      

      Dependencies and security

      Dependency updates

      Dependency name | Previous version | New version
      ---|---|---
      lit | 3.2.1 | 3.3.3
      receptor | 1.0.0 | --

      Note: receptor has been removed as a dependency. Its functionality has been reimplemented in first-party code.
      Thanks @aduth! (#6489)

      Dev dependency updates

      Dependency name | Previous version | New version
      ---|---|---
      @babel/core | 7.26.8 | 7.29.7
      @babel/preset-env | 7.26.8 | 7.29.7
      @chanzuckerberg/axe-storybook-testing | 6.3.1 | --
      @material-design-icons/svg | 0.14.13 | 0.14.15
      @rollup/plugin-commonjs | 28.0.3 | 29.0.3
      @spiriit/vite-plugin-svg-spritemap | 4.0.0 | 6.0.0
      @storybook/addon-a11y | 6.5.16 | 9.1.20
      @storybook/addon-essentials | 6.5.16 | --
      @storybook/addon-links | 6.5.16 | --
      @storybook/builder-webpack5 | 6.5.16 | --
      @storybook/html | 6.5.16 | --
      @storybook/html-vite | -- | 9.1.20
      @storybook/manager-webpack5 | 6.5.16 | --
      @storybook/test-runner | -- | 0.23.0
      @types/node | 20.14.10 | 24.13.3
      @uswds/compile | -- | 1.3.2
      autoprefixer | 10.4.20 | 10.5.0
      axe-core | 4.10.2 | --
      axe-playwright | -- | 2.2.2
      concurrently | -- | 10.0.3
      css-loader | 6.8.1 | --
      del | 6.0.0 | 8.0.1
      esbuild | -- | 0.28.1
      eslint | 8.56.0 | 10.8.0
      eslint-config-airbnb-base | 15.0.0 | --
      eslint-config-prettier | 9.1.0 | 10.1.8
      eslint-plugin-airbnb-base | 0.0.1-security | --
      eslint-plugin-import | 2.31.0 | --
      eslint-plugin-import-x | -- | 4.17.1
      eslint-plugin-lit | 2.0.0 | 2.3.1
      eslint-plugin-no-unsanitized | 4.1.2 | --
      file-loader | 6.2.0 | --
      globals | -- | 17.8.0
      gulp | 4.0.2 | 5.0.1
      gulp-mocha | 9.0.0 | 10.0.1
      gulp-postcss | 9.0.1 | 10.0.0
      gulp-rename | 2.0.0 | 2.1.0
      gulp-sass | 6.0.0 | 6.0.1
      html-webpack-plugin | 5.6.3 | 5.6.8
      http-server | -- | 14.1.1
      magic-string | -- | 0.30.21
      merge-stream | 2.0.0 | --
      mocha | 10.8.2 | 11.8.0
      postcss | 8.5.2 | 8.5.25
      postcss-discard-comments | 6.0.2 | 8.0.2
      postcss-import | 15.1.0 | --
      postcss-loader | 7.3.3 | --
      postcss-preset-env | 9.6.0 | --
      prettier | 3.4.2 | 3.9.6
      react-dom | 17.0.2 | --
      resolve-url-loader | 5.0.0 | --
      sass-embedded | 1.83.4 | 1.100.0
      sass-loader | 16.0.4 | --
      sass-true | 6.0.1 | 10.1.0
      sinon | 12.0.1 | 22.1.0
      snyk | 1.1295.3 | 1.1306.2
      storybook | -- | 9.1.20
      style-loader | 3.3.3 | --
      svgo | 3.3.2 | 4.0.2
      twig | -- | 3.0.0
      twigjs-loader | 1.0.3 | --
      vite | 6.2.2 | 6.4.3
      vite-plugin-svg-sprite | 0.6.2 | --
      wait-on | -- | 9.1.0
      webpack | 5.98.0 | 5.109.2
      webpack-cli | 5.1.4 | 7.2.2

      0 vulnerabilities in regular dependencies (dependencies for USWDS projects installed with npm install @uswds/uswds)

      19 vulnerabilities (11 moderate, 3 high) in devDependencies (development dependencies)

      SHA-256 for release

      da91c65e6fc736fa397f0daf6ca2c2c95711506d85c7ad2537a4570725401e1c

      Additional contributions

    4. 🔗 exe.dev Have an Agent Babysit Your Deployments rss

      Deployments are scary. That’s the moment you break things.

      Not-deployments are even scarier. Waiting just makes the next deployment bigger.

      As the saying goes: “If it hurts, do it more.”

      The obvious, correct answer is CD. But then you’re off building canaries and waves and automated detection systems as gates. Canaries and waves are easy. Automated detection systems are hard. There are an indefinite number of things that can go wrong, and missing one of them takes you down. The asymmetry there is exactly the same asymmetry that makes deployments scary in the first place.

      The historical answer was: It’s a lot of engineering effort and a lot of pain. And so CD gets delayed, and humans babysit deployments, and deployments happen infrequently, and the cycle of inefficient misery and fear continues.

      This has exactly the right shape for an agent instead of code: Lots of rich data, a very long tail of possible states, relatively few runs (a handful a day, not 100qps).

      And we now have intelligence on tap. Let’s use it!

      At exe, Athena oversees our deployments. (All our bots have names, but that’s just so it’s easy to talk about them. They’re programs, not people.)

      Athena sits in a system called “exe-ops” which is our Deployment Command Center. We started with shell scripts, but then built a UI. Traditionally, you do a migration to something like Spinnaker. Instead, we’re building up from shell scripts into the exact shape we want. Athena is part of that story.

      The bot has read access to git and metrics and logs. It decides at each stage: should we proceed? Which machines should be in the next wave? It can escalate to a human and it can pause a deployment—or refuse to start one, if it deems it unwise. It communicates by sending us Slack messages.

      It’s great! It is diligent and thorough. It reads the diffs, analyzes the logs, checks for unforeseen issues, self-heals around weird problems, and reports on how to make future runs smoother.

      I could probably oversee deployments better than Athena. But the important question is not “in theory, could I do a better job?” but “in reality, will I do a better job?” We’re all busy. Athena does a much, much better job than I actually would.

      Athena lets me focus my attention elsewhere, until something happens that’s worth my intervention. And by deploying more often, those interventions are rarer and smaller.

    5. 🔗 tintinweb/pi-subagents v0.17.1 release

      No content.

    6. 🔗 r/LocalLLaMA Memory prices climb 500% in 12 months, up to 10x the lowest ever tracked prices - 128GB of DDR5 now $3,399 rss
    7. 🔗 HexRaysSA/plugin-repository commits sync repo: +2 releases rss
      sync repo: +2 releases
      
      ## New releases
      - [diaphora](https://github.com/joxeankoret/diaphora): 3.4.1
      - [eject_idb](https://github.com/allthingsida/eject_idb): 0.0.4
      
    8. 🔗 3Blue1Brown (YouTube) The jumping pegs puzzle rss

      Part of a series of monthly puzzles with MoMath.

    9. 🔗 @binaryninja@infosec.exchange What do Binary Ninja workflows do for you? A lot! Check out what mastodon

      What do Binary Ninja workflows do for you? A lot! Check out what @mei managed to pull off using them to clean up conditional jump threading:

      https://codeberg.org/mei-b/bn-analysis- improvements

    10. 🔗 r/LocalLLaMA Running DeepSeek V4 Flash Q4_K_XL at ~100 tok/s prompt processing on 4× RTX 3060 12GB rss

      Running DeepSeek V4 Flash Q4_K_XL at ~100 tok/s prompt processing on 4× RTX 3060 12GB | I managed to run the 143–144 GiB DeepSeek-V4-Flash-0731 UD-Q4_K_XL GGUF on four RTX 3060 12GB cards while keeping a 360k–376k context window. Hardware: CPU: Intel Core i9-10920X, 12C/24T RAM: 128 GB DDR4-3200, quad-channel GPU: 4× NVIDIA RTX 3060 12GB Total VRAM: 48 GB Storage: NVMe SSD Engine: llama.cpp, build b10181 Model: unsloth/DeepSeek-V4-Flash-0731-GGUF Quant: UD-Q4_K_XL, approximately 144 GiB KV cache: Q8_0 The best high-speed configuration so far: llama-server \ -m DeepSeek-V4-Flash-0731-UD-Q4_K_XL-00001-of-00005.gguf \ -c 368640 \ -ncmoe 34 \ -ts 100,1,1,1 \ -ot 'blk.(3[4-6]).ffn_.exps=CUDA1,blk.(3[7-9]).ffn.exps=CUDA2,blk.(4[0-2]).ffn.*_exps=CUDA3' \ -ctk q8_0 \ -ctv q8_0 \ -b 2048 \ -ub 2048 \ -np 1 \ -lm none \ --threads 20 \ --flash-attn on Measured with a roughly 20.5k-token prompt: Configured context: 368,640 tokens Prompt processing: 99.4 tok/s Text generation: 10.1 tok/s Minimum free VRAM under load: GPU0: 671 MiB GPU1: 842 MiB GPU2: 1395 MiB GPU3: 1395 MiB Model load time: approximately 198 seconds Other measured context/safety options: Context Prefill Decode Minimum free VRAM 376832 99.5 t/s 10.4 t/s 611 MiB 368640 99.4 t/s 10.1 t/s 671 MiB 360448 99.4 t/s 10.1 t/s 735 MiB The interesting part is the GPU layout. -ncmoe 34 keeps the experts from blocks 0–33 in system RAM. The remaining nine expert layers are explicitly distributed across GPUs 1–3, three layers per GPU. The extreme -ts 100,1,1,1 split does not distribute those explicitly assigned expert weights. Instead, it pushes most non-expert tensors—attention, KV-related allocations, etc.—onto GPU0. That leaves enough space on GPUs 1–3 for the large expert layers. This was much better than trying to calculate the layout analytically. With -ncmoe and explicit -ot overrides, tensor placement is discrete and somewhat unintuitive, so I measured every candidate. Microbatch size was the biggest performance lever: -ub 1024: approximately 63.4 tok/s prompt processing -ub 2048: approximately 99.4 tok/s prompt processing Decode remained almost unchanged at approximately 10.1–10.5 tok/s. At the full 393,216-token context, -ub 2048 also worked, but GPU0 had only 493 MiB free under load. Reducing the configured context to 368,640 restored a 671 MiB margin without reducing prompt-processing speed. For comparison, the safer -ub 1024 configuration can run with a configured context of 524,288 and still showed about 1032 MiB free on the tightest GPU, but prompt processing drops to approximately 63.4 tok/s. A few additional findings: Q8_0 KV is the default choice. F16 KV at c=393216 left only 587 MiB free. -ncmoe 33 caused a CUDA allocation failure. Memory mapping was disabled with -lm none. -np 1 is important; multiple slots multiply KV-cache requirements. The model is mostly in system RAM, so quad-channel memory bandwidth matters heavily. Even so, getting approximately 100 tok/s prompt ingestion and 10 tok/s generation from a 144 GiB MoE model on four consumer 12GB GPUs is much better than I expected. The configuration has been tested under real prompt load. The entire 368k context window has not yet been filled end-to-end, so the number above is the configured capacity, not a claim that I already completed a 368k-token generation test. Generated by ChatGPT 😂. submitted by /u/syscomua
      [link] [comments]
      ---|---

    11. 🔗 tintinweb/pi-subagents v0.17.0 release

      ⚠️ Note — an agent file's frontmattername: now substitutes for the filename as its subagent_type. Following Claude Code, the declared name is the dispatch identity and the filename is only the fallback, so blubb.md with name: code-review is spawned, mentioned and listed as code-review. Any value is accepted except one containing :, which Claude Code reserves for plugin scoping; such a file — like any unparseable one — is skipped with a warning rather than loaded under a name nothing honours, and strictAgentFiles turns that into a startup failure.

      ⚠️ Breaking — subagent sessions now persist to disk by default (rememberAgents). Transcripts that used to be in-memory are written to the session dir and appear nested under their spawner in pi's /resume. Per- agent persist_session: overrides it in both directions — false keeps that agent in memory, true persists it even with the setting off. Set rememberAgents: false to restore the previous behaviour globally.

      pi-subagents-mention.mp4

    12. 🔗 r/LocalLLaMA Qwen dev says not to wait for 35B-A3B rss

      Qwen dev says not to wait for 35B-A3B | What does this mean? Is there something else coming? Maybe 122B? Or no models? submitted by /u/Mean-Ad1493
      [link] [comments]
      ---|---

    13. 🔗 Ampcode News Education Discount rss

      Students and teachers can now subscribe to Amp for $10/month, half the usual price.

      What do you get for $10? Quite a lot!

      You get to use the best frontier agent. You get orbs, our remote machines that let you run agents from anywhere without supervision.

      You get code hosting for unlimited public/private repositories.

      And you get to use great models, through linking your ChatGPT sub for GPT-5.6 or 𝕏 Premium+/SuperGrok subscription for Grok 4.6. Plus $10 in credits each month for use on any other model.

      Get it at ampcode.com/edu.

    14. 🔗 Ampcode News Talk to Puck rss

      You can now talk with Puck in realtime:

      Realtime chat with Puck is powered by gpt-realtime-2.1. It delegates work to the Puck agent you already know, powered by GPT-5.6 Sol. Once the agent responds, gpt-realtime-2.1 summarizes the answer out loud while the full response appears in the thread.

      You can now have a proper back-and-forth conversation with Puck without waiting for text responses to stream in. Just talk. That makes it easier to coordinate parallel work, get progress updates, send follow-up instructions, or talk through big ideas without typing, whether you are sitting in front of your computer or on the go.

      Here are some conversation starters we have used:

      • "Check my active threads and tell me which ones need input."
      • "Review the messages in the #issues Slack channel where I am tagged. Read them to me one by one so we can talk through how to fix each one."
      • "Start an agent to fix the CI failure, tell me what caused it, and keep me updated on the fix."
      • "The executor lease reconciliation pipeline broke. Tell me about the changes made to it yesterday."
  3. August 17, 2026
    1. 🔗 IDA Plugin Updates IDA Plugin Updates on 2026-08-17 rss

      IDA Plugin Updates on 2026-08-17

      New Releases:

      Activity:

      • capa
        • b3817357: build(deps-dev): bump pyinstaller from 6.21.0 to 6.22.0 (#3157)
        • 84e36402: Sync capa-testfiles submodule
        • 1218e3b7: build(deps): bump setuptools from 83.0.0 to 84.0.0 (#3156)
      • chernobog
        • d272b5df: fix: stub HEXDSP in standalone tests to prevent crash on AST destruction
        • 79f1488d: ci: separate cache restore and save steps
        • 34c4f543: feat: expose full capability surface through IDC interpreter
      • disrobe
        • a5aa5b8e: recovery: expand appimage, flutter, php, jvm, and javascript
        • 65755cfa: recovery: recover python try loops, go control edges, and aarch64 ari…
        • d6d96518: javascript: recover system register parameter names
        • 4c01ca15: jvm: recover kotlin nested finally copies
        • 1471a1c1: query: expose canonical instruction effect rows
        • 92e5a255: javascript: recover rollup iife parameter names
        • 33e85455: recovery: expand erofs, go and php with bounded pdb argument lists
        • 8fa80485: recovery: expand erofs, nativeaot, flutter, python, php, wasm and d
        • a7a9fec7: recovery: expand firmware, python, javascript and witness output
        • 0e7c8980: recovery: add firmware, crc32 and report redaction
        • 81d68e0d: recovery: expand native, dalvik, php, as3 and wasm output
      • ffxiv_bossmod
      • ida-pro-mcp
      • plugin-ida
        • 0fe1e8c2: chore(deps): Bump step-security/harden-runner from 2.20.1 to 2.21.0 (…
      • project
        • fd1bb70b: updated webapppapebapp web notifications
      • twdll
        • 4a97c0e8: chore: update gitignore
        • 1837efd4: feat: add encyclopedia url func
        • 1dafd620: refactor: new database api
        • 02f96bbd: tests: improve set party test
        • e5abe584: refactor: structs refactor
        • 0a195e6b: chore: add manual test template
        • 3c7d1e70: feat: add SetPrimaryParty & SetParty for char
    2. 🔗 PrimeIntellect-ai/prime-agent v0.7.3 release
      • Fixed assistant rendering when provider payloads contain null or sparse content blocks.
      • Added authenticated host-request contracts with per-call request IDs, generation fencing, cancellation signals, and currentness checks.
      • Fixed root daemon shutdown retaining cleanup ownership while kill events are in flight.
      • Changed RLM family discovery to use a daemon-owned append-only spawn ledger with per-child display metadata instead of reconstructing topology from session files.
      • Fixed long-running macOS supervisors losing ownership when system cleanup removed authority records from $TMPDIR.
      • Fixed deleted RLM children leaking kernel snapshots while retaining their readable transcript tombstones.
      • Changed Agents View subagent rows to show stable name · model/effort · summary metadata.
      • Changed the default Cerebras model to the available gpt-oss-120b route and aligned cross-provider handoff fixtures with the generated catalog.
      • Fixed the agent going silent after an automatic context compaction interrupted unfinished work: the tool loop now resumes when a threshold compaction fails or is skipped, and active goals keep continuing after a successful mid-goal threshold compaction.
      • Changed the agents view splash hint from "type to start" to "type to search sessions".
      • Added app.edits.expand (ctrl+j) to toggle edit diffs; diffs are now shown only by this toggle, and ctrl+o no longer affects them.
      • Changed edit rendering so the ╰─ <path> +N -M summary line is always visible and ctrl+j toggles the diff inline beneath it, indented to the summary text.
      • Fixed fullscreen wheel scrolling in Ghostty while retaining application link clicks; set terminal.fullscreenMouse to false to use native Cmd-click instead.
      • Changed the agents view to sort idle and inactive sessions by last message time, newest first, while keeping running agents in stable creation order.
      • Fixed openai-codex models being invisible to rlm subagents and find_models because model discovery reported Prime Agent's own version as the Codex client version (#1375 by @bilelrais).
      • Added a working hint that recommends sharing traces with Prime Intellect to help train open-source LLMs.
      • Restored bare prime-agent --resume opening the agents view and the /resume [id|path] slash command; bare commands open the agents view and an argument resumes that session in place.
      • Fixed URLs not opening on click in fullscreen mode on terminals such as Ghostty; clicking a link in the transcript, dock, or overlays now opens it in the browser.
      • Fixed ctrl+p ("Toggle agent message expansion") only toggling received agent messages; it now expands and collapses sent agent messages together with received ones.
    3. 🔗 anthropics/claude-code v2.1.234 release

      What's changed

      • Added the optional CLAUDE_CODE_PROJECT_DIR_NAME environment variable: hosts that give each session its own config directory can choose a short name for the per-project transcript directory
      • Added the selection:clear keybinding action, so a key can be bound to clear an in-app text selection; also works in the agents view
      • Added a GitLab merge request badge to the footer and statusline: repos with a GitLab remote and an authenticated glab CLI show MR !N with draft/pending/green states
      • Claude Code now continues your session automatically when a claude.ai usage limit resets; turn it off in /config ("Continue automatically at usage limit")
      • Claude is now told to use your account email only to identify you, and not to send it to unrelated services unless you ask
      • Security: remote file reads, session restore, CLAUDE.md includes, workflow scripts and file uploads now reject Windows NT-namespace (\??\) paths, hardening the remaining pre-approval file accesses against the NTLM credential-leak vector
      • Fixed auto mode in very long sessions repeatedly re-checking and denying sandboxed commands' network access after the conversation had been compacted
      • Fixed session-scoped permission answers (including denies) being dropped when answering background subagent tool permission prompts
      • Fixed a crash when an API response on the non-streaming fallback path (typically via third-party gateways) contained a thinking block missing its thinking field or a text block missing its text field
      • Fixed markdown rendering becoming extremely slow for some messages containing unusual Unicode sequences
      • Fixed SendMessage rejecting a recipient copied from ListAgents when the session name is at the 200-character cap or emoji-heavy
      • Fixed repository detection mis-reading the host of git remotes with unusual userinfo, producing links and repo-specific behavior for the wrong host
      • Fixed MCP diagnostics printing resolved secrets: scope-conflict warnings now show the configured ${VAR} form, and connection-failure details show only the server origin
      • Fixed strictKnownMarketplaces allowlists accepting SCP-style git marketplace sources whose host differs from the one git would actually connect to
      • Fixed modal text such as the /login OAuth URL losing characters when copied in fullscreen
      • Fixed a --- horizontal rule in rendered markdown running into the line after it
      • Fixed consecutive shell commands splitting into multiple "Ran 1 shell command" rows when todo/task updates were interleaved between them
      • Fixed dialogs like /permissions opened while a ! shell command was running being dismissed when the command finished
      • Fixed a queued ! shell command being sent to the model as plain text after pressing up-arrow to edit the queued input
      • Fixed queued messages reappearing in the prompt history while still queued, Esc while selecting a queued message no longer interrupts the turn, and ! mode no longer sticks after a mid-turn submit
      • Fixed accepting the "Try the new fullscreen renderer?" prompt restarting the session without its permission mode (e.g. --dangerously-skip-permissions), tool allow/deny rules, model or effort flags
      • Fixed /tui dropping launch --allowed-tools/--disallowed-tools rules when it restarts; it now declines to switch, with the reason, when the session has restrictions a restart can't carry over
      • Fixed trust prompts omitting the repository-wide scope warning when the directory was first seen before the repository existed there
      • Fixed a case where an IDE diff tab closing during a permission re-prompt could answer the new prompt with the previous input
      • Fixed: files sent to the user during Remote Control sessions hosted by Claude Code Desktop or VS Code now upload, so they open on phone and web instead of showing an empty card
      • Fixed: after /login while CLAUDE_CODE_OAUTH_TOKEN is set, the stale-token reminder no longer leaks into Claude's automatically resumed turn — it now appears only to you
      • Fixed: permission previews now relay only to channel servers admitted by the inbound trust gate, and a server's explicit permission-capability opt-out is honored
      • Fixed: credential masking on relayed permission previews can no longer hide commands, paths, or destinations from the approver; oversized private-key blocks now redact under full-strength redaction
      • Fixed: provider API tokens that mask on permission previews now mask even when directly followed by shell delimiters
      • Fixed Claude Desktop inter-session messages being silently dropped by the recipient session when cross-session messaging read as disabled, which left the sender's query "thinking" for many minutes
      • Remote Control: signing this computer in to a different claude.ai account or organization now stops the running session within seconds and says why, instead of a misleading HTTP 404 hours later
      • Remote Control sessions started from Claude Code Desktop or VS Code now keep phones and claude.ai/code updated on the session's permission mode (and claude.ai/code on the model) as they change
      • Remote Control: effort picks made on a phone or on claude.ai/code now apply to terminal- and Desktop/VS Code-hosted sessions, and the session publishes its effort level to connected clients
      • SendMessage and ListAgents now say when your account's session list was too long to check completely, instead of treating unseen sessions as absent
      • Expired Anthropic profile credential now points you at /login when a claude.ai login would take precedence
      • Improved the transcript: your own prompts now render markdown (highlighted code blocks, inline code, lists) the same way replies do
      • Improved the "API returned an empty or malformed response" error to say what came back (content type, body kind, size, request ID) and why the original streaming request failed
      • Improved auto-generated session titles to read as short, specific names (e.g. "Login button bug") rather than sentences restating your request (e.g. "Fix the login button on mobile")
      • Reduced the context cost of loading the built-in claude-api skill from ~200k+ tokens to ~25k by loading reference docs on demand
      • /permissions can now be opened while Claude is working — rule changes apply to the rest of the current turn
      • /add-dir <path> can now be used while Claude is working; /add-dir, /autocompact, /theme, /help, /config and /advisor dialogs open mid-turn in the fullscreen TUI
      • /goal now clears itself with a notice when a turn dies on an unrecoverable error (e.g. revoked auth, an exhausted credit balance, or a context overflow) instead of staying armed
      • /goal: when background tasks keep a goal waiting for 30+ minutes, Claude now checks in on them instead of waiting indefinitely (set CLAUDE_CODE_GOAL_CHECKIN_MINUTES=0 to opt out)
      • claude setup-token now rejects unexpected extra arguments instead of silently ignoring them
      • Changed Esc in fullscreen mode to no longer clear a mouse text selection: it interrupts or dismisses as usual and the selection stays highlighted
      • Removed the redundant "Allowed by auto mode classifier" line that auto mode showed under every Agent tool call
      • Removed the "Default teammate model" setting from /config; agent-team teammates now use the leader's model unless the spawn names one
      • Dimmed the elapsed-time counter on the running tool header so it no longer competes with the bold counts
      • Background task notifications delivered between turns are now sent to the model inside <system-reminder> tags, matching mid-turn delivery
      • Mantle: skip the admin-pin availability probe at startup when a main-loop model is already picked
      • Windows: startup no longer stalls on repeated rename retries when ~/.claude.json is read-only
    4. 🔗 backnotprop/plannotator v0.27.4 release

      Follow @plannotator on X for updates

      Missed recent releases? Release | Highlights
      ---|---
      v0.27.3 | Folder watcher freeze fix on large repos, first SBOM-attested release pipeline
      v0.27.2 | Mobile plan and code review, Codex CLI 0.147 fix, folder annotate cold-start, configurable markdown extensions
      v0.27.1 | Open-in-editor launch fix, file headers respect Viewed/Git-add visibility toggles
      v0.27.0 | Call Flow analysis, --tailscale remote reviews, review panel remembers your view, Pi rebuild (breaking command rename), focus-mode shortcut
      v0.26.8 | Placed comment markers on HTML pages, shift-click multi-select, live app annotation
      v0.26.7 | Pinpoint targets any element on HTML pages, smarter hover labels, zero-scan hit testing
      v0.26.6 | Fixed empty environment variables in sandboxed sessions (Bun 1.3.14 builds)
      v0.26.5 | HTML pinpoint element annotations, durable annotate submissions, installer fallback for old git, vim HUD cursor fix
      v0.26.4 | Skill-menu hover jitter fix (same-day patch on v0.26.3)
      v0.26.3 | Skill references in comments with / or $, reachable remote session URLs, worktree switcher tooltips
      v0.26.2 | Single-file diff tabs render fully, no more silently dropped review files, light/dark theme pairs, palette-matched code blocks
      v0.26.1 | GitButler 0.22.0 compatibility via capability-probed JSON flags

      What's New in v0.27.4

      A Guided Review can now leave Plannotator. This release ships portable guide exports, share links on guides.show, and a guide CLI any agent can drive, alongside a favicon style switcher, jj support for Call Flow, GitLab artifact fixes in PR review, and a smoother call-flow Lens. Eighteen PRs, four from community contributors, two of them first-timers.

      Portable Guided Reviews and guides.show

      Guided Reviews used to live and die inside your review session. Now a guide has three ways out:

      Download it. Every guide gets a "Download portable guide" button that produces one HTML file containing the full guide and the diff it describes. It opens anywhere, renders exactly like the in-app guide with side-by-side diffs and per-section reviewed checkboxes, and needs no Plannotator install. The file stays small because it carries your content, not the renderer: the viewer loads from guides.show, pinned by filename and cryptographic checksum, so a tampered or wrong viewer never executes. Offline, the file degrades to a readable plain-text version of the guide.

      Share it. "Create share link" uploads the guide to guides.show and hands you a link anyone can open in a browser. Shares are end-to-end encrypted by default: the key lives in the URL fragment after the #, which browsers never send to the server, so guides.show stores bytes it cannot read. You also get a one-time delete token, and "Remove link" works from the same dialog for as long as that Plannotator remembers the share. An optional "Allow link previews" checkbox stores the guide unencrypted so chat apps can show its title; that is a choice, never the default. Setting PLANNOTATOR_SHARE=disabled turns all of this off.

      Author it from anywhere. The new plannotator guide subcommands (list, export, share, unshare) let any agent or script produce and publish a guide from a guide JSON and a patch, without a browser in the loop.

      Saved guides from v0.27.x load unchanged. The share service runs on Cloudflare with add-only, content-hashed viewer publishing and per-IP rate limiting on creation.

      Choose your favicon: Totman or the classic P

      The browser-tab icon is now a setting. Appearance settings offer two styles with visual previews: Totman, the current mascot, and Classic P, the original Plannotator mark restored byte-for-byte from the pre-mascot era. The server remembers your choice and serves it directly, so tabs show the right icon from the first paint without flashing the default. Hosts that embed the published UI packages are unaffected unless they opt in.

      Call Flow analysis on jj repositories

      Call Flow previously required a plain Git checkout. Reviews running on jj (Jujutsu) colocated repos now get the same changed-call-path analysis: the jj snapshot is resolved to the underlying Git objects and fed to the same CallDiff engine, with the same per-file Lens and dock views. Diff collection is untouched; this only extends where the analysis can run.

      GitLab PR artifacts fetch reliably and more safely

      Reviewing GitLab merge requests with uploaded artifacts (screenshots, logs, design files) got a hardening pass. Uploads now fetch through the authenticated API with a strict rewrite that only touches real upload URLs, falls back to the original web route when a self-hosted GitLab predates the API route, maps 401/403 responses to a clear "run glab auth login" hint, and no longer serves HTML or JavaScript content types through the artifact proxy. A regression test pins the invariant that credentials never follow a cross- origin redirect.

      The call-flow Lens stops fighting your scroll

      Community feedback within hours of trying Call Flow in Safari: the per-file Lens popover closed randomly mid-scroll and popped open for every badge that passed under the cursor. Three causes, three fixes: the Lens's internal scroll no longer chains to the page when momentum hits its edge (the chain moved the popup out from under a stationary pointer, which read as a random close and was worst under Safari rubber-banding); hover now has a 100ms intent delay so drive-by badges stay closed; and an in-flight page scroll holds any pending close until the scroll settles.

      Reported by Rustan (@acewhocares on X).

      Additional Changes

      • Touch selection survives the comment composer. On phones and tablets, dragging a multi-line range in a single-file diff no longer collapses the selection when the composer opens; the range you dragged is the range you comment on. #1333
      • Skill picker works with screen readers. The / and $ skill reference menu now exposes real listbox semantics with option roles and active-descendant tracking, so assistive tech announces what Enter will insert, closing #1233. #1316 by @ashish921998
      • Blog: an interactive UI for the grill-me skill. A new post on using /plannotator-last as the review surface for Matt Pocock's grill-me workflow, at plannotator.ai. #1321, #1322, #1323, #1332
      • Security page linked from the site footer. #1305

      Install / Update

      macOS / Linux:

      curl -fsSL https://plannotator.ai/install.sh | bash
      

      Windows:

      irm https://plannotator.ai/install.ps1 | iex
      

      Claude Code Plugin: Run /plugin in Claude Code, find plannotator , and click "Update now".

      OpenCode: Clear cache and restart:

      rm -rf ~/.bun/install/cache/@plannotator
      

      What's Changed

      • fix(comments): expose skill picker semantics to assistive tech by @ashish921998 in #1316
      • feat(review): jj support for Call Flow analysis by @graemefolk in #1312
      • blog: an interactive UI for the grill-me skill by @backnotprop in #1321
      • blog: grill-me post additions by @backnotprop in #1322
      • blog: repo link, image alt text, and larger blog type by @backnotprop in #1323
      • feat: Portable Guided Reviews, export, share links, agent-authored guides, guides.show by @backnotprop in #1324
      • guides-show: GitHub link in the landing page header by @backnotprop in #1327
      • guide-viewer: label agent harnesses in the generated-by line by @backnotprop in #1328
      • guide-viewer: readable on phones and tablets, desktop untouched by @backnotprop in #1329
      • seo: index live root blog pages by @backnotprop in #1332
      • guide: voice rules in the organizer prompt by @backnotprop in #1330
      • docs(marketing): link security page from footer by @backnotprop in #1305
      • fix(review): preserve dragged diff ranges on compact touch before commenting by @backnotprop in #1333
      • feat(ui): Totman/Classic P favicon style switcher by @FNDEVVE in #1325
      • fix(review): GitLab upload artifact fetching via authenticated API with hardened rewrite by @yuensunn in #1228
      • guides-show: example guide screenshot at the bottom of the landing page by @backnotprop in #1336
      • guides-show: example guide screenshot replaces the abstract figure, opens in a lightbox by @backnotprop in #1337
      • fix(review): stop the call-flow Lens closing mid-scroll and opening on drive-by hovers by @backnotprop in #1338

      New Contributors

      Contributors

      Four community authors shipped code in this release, two for the first time:

      • @FNDEVVE built the favicon style switcher in #1325, including restoring the classic P icon exactly as it shipped before the mascot era, and worked through a review round that added server-side icon serving so the choice applies without a flash. Their second contribution.
      • @graemefolk extended Call Flow analysis to jj repositories in #1312, their third contribution to Plannotator's jj support, which they have carried since the original provider landed.
      • @yuensunn fixed GitLab merge request artifacts in #1228, their first contribution, and stuck with it through a security-focused review round on the URL rewrite.
      • @ashish921998 made the skill reference menu real for screen reader users in #1316, their first contribution.
      • Rustan (@acewhocares on X) test-drove Call Flow in Safari and reported the Lens scroll behavior that #1338 fixes, hours after trying the feature.

      Full Changelog : v0.27.3...v0.27.4

    5. 🔗 r/LocalLLaMA Artificial Analysis' Qwen3.8-27B benchmarks put it neck and neck with DeepSeek V4 and GPT-5.6 Luna Max rss
    6. 🔗 @binaryninja@infosec.exchange Seven sidebars was a bit much. Sidekick 26.1 brings Indexes, Notebooks, Code mastodon

      Seven sidebars was a bit much. Sidekick 26.1 brings Indexes, Notebooks, Code Maps, and Repositories together in one Sidekick Resources sidebar! Search across all four from the same place, then open or pin whatever you need right in Binary Ninja. See what else is new in 26.1: https://sidekick.binary.ninja/blog/sidekick-26-1-a-proper-home-for- sidekick/#seven-sidebars-were-too- many

    7. 🔗 @malcat@infosec.exchange Did you know that [#Kesakode](https://infosec.exchange/tags/Kesakode) can use mastodon

      Did you know that #Kesakode can use fuzzy-matching for functions? While a bit slower, this kind of lookup helps against obfuscation (here: #vidar)

    8. 🔗 r/LocalLLaMA After pushing 1M+ tokens through Qwen 3.8 27B, here is my optimal llama.cpp config for 16GB VRAM (73k Context, Agentic Coding) rss

      After pushing 1M+ tokens through Qwen 3.8 27B, here is my optimal llama.cpp config for 16GB VRAM (73k Context, Agentic Coding) | Dando seguimiento a mi post anterior sobre cómo tengo montado mi servidor de presupuesto (Intel N100 + RTX 5060 Ti 16GB), varios me preguntaron por una mirada más profunda a mi configuración real de inferencia y al desempeño agentic en el mundo real. Como muchos de ustedes, estaba refrescando la página esperando descargar Qwen 3.8 27B apenas salió. Después de pasar todo el fin de semana estresándolo con flujos de trabajo de codificación agentic, logré correr un proyecto completo y grande casi todo de forma autónoma (más de 1M de tokens procesados en total , solo 3 prompts). Aquí va un resumen rápido de la configuración base antes de meternos en los detalles del config y del workflow.

      Specs y parámetros rápidos

      • Modelo: Qwen3.8-27B-UD-Q3_K_XL.gguf
      • Hardware: RTX 5060 Ti (16GB VRAM) + Intel N100 (4C/4T, 16GB RAM)
      • Ventana de contexto: 73,728 (73k de contexto) corriendo tranqui en 16GB de VRAM.
      • Cuantización de KV Cache: q4_1 para el contexto principal
      • Decodificación especulativa: MTP nativa activada (spec-type = draft-mtp, n-max = 2)
      • Sampling: temp = 0.65, top_p = 0.95, top_k = 20, min_p = 0.05

      El experimento: armar una API completa con 3 prompts

      En vez de correr benchmarks sintéticos, metí esta configuración por una cadena real de ingeniería de software: construyendo una REST API no oficial y un servidor MCP para un foro vBulletin heredado.

      1. Prompt 1 (Arquitectura del sitio y análisis): Pedí al modelo que mapee el sitio objetivo. Generó una especificación en Markdown impecable de ~1,500 líneas que cubría análisis estructural, nodos HTML rescatables, payloads JSON esperados, selección de stack, lógica de paginación, autenticación de sesión y endpoints de búsqueda—mucho más a fondo de lo que yo habría escrito a mano.
      2. Prompt 2 (Arquitectura de desarrollo): Usando la spec como única fuente de verdad, diseñó un plan de implementación modular de NestJS dividido en 9 fases de ejecución:
      3. Fase 1: Estructura inicial del proyecto
      4. Fase 2: Modelos de dominio
      5. Fase 3: Scraping core (HTTP + limitación de tasa + reintentos)
      6. Fase 4: Parsers de HTML (cheerio)
      7. Fase 5: Capa de caché
      8. Fase 6: Servicios de aplicación + REST API
      9. Fase 7: Autenticación (sesiones con cookies)
      10. Fase 8: Servidor MCP (entrega principal)
      11. Fase 9: Fortalecimiento, documentación y entrega
      12. Prompt 3 (Ejecución autónoma agentic): La prueba de verdad. Le pedí a OpenCode (usando Qwen 3.8 27B) que actuara estrictamente como orquestador, creando sub-agentes para cada fase de tareas. Corrió de forma autónoma por ~2 horas. Cuando se acercaron los límites de contexto, OpenCode resumió su estado y siguió construyendo. Escribió tests unitarios, aplicó linting y entregó código 100% funcional—solo necesitando un arreglo automatizado menor cuando le di un payload de HTML crudo con un caso extremo.

      El archivo de configuración llama.cpp

      Aquí está mi archivo exacto de configuración de enrutador --models-preset . Fíjate cómo fit = off se usa en el perfil de 27B junto con ctx-size = 73728 (73k) y q4_1 para cuantizar la KV cache, con el objetivo de maximizar la asignación de VRAM mientras se mantiene el rendimiento nativo de MTP. ```ini

      ==============================================================================

      LLAMA.CPP — CONFIGURACIÓN DE INFERENCIA (modo router / --models-preset)

      ==============================================================================

      Objetivo de hardware:

      GPU: 16 GB VRAM (RTX 5060 Ti)

      CPU: Intel N100, 4C/4T (Debian Headless)

      ------------------------------------------------------------------------------

      GLOBAL / LÍNEA BASE

      ------------------------------------------------------------------------------

      [*]

      --- HILOS DE CPU


      Reserva 1 core para SO/servicios durante el decode.

      Usa los 4 threads durante ráfagas de prefill del prompt.

      threads = 3 threads-batch = 4

      --- SERVIDOR / CONCURRENCIA


      Un solo slot; desactivado continuous batching para máximo rendimiento por

      usuario.

      parallel = 1 cont-batching = 0

      --- GPU / AJUSTE DE VRAM


      flash-attn = on fit = on

      Holgura de seguridad para el límite físico de VRAM (MiB).

      Ponlo bajo (128) porque el sistema es headless (100% VRAM disponible para

      inferencia).

      NOTA: Si usas caches KV draft de MTP, ojo con la asignación doble de VRAM.

      Sube a 128-256 si te topas con OOMs.

      fit-target = 128

      --- CONTEXTO & CACHÉ ------------------------------------------------------

      ctx-size = 65536 context-shift = 1

      Desactiva checkpoints de contexto (evita problemas de reprocesamiento en

      arquitecturas híbridas)

      ctx-checkpoints = 0

      RAM Prompt Cache (2 GiB)

      cache-ram = 2048

      --- KV CACHE GLOBAL


      cache-type-k = q5_1 cache-type-v = q5_1

      --- PREFILL / BATCHING


      batch-size = 2048 ubatch-size = 1024

      --- SAMPLING POR DEFECTO (Códigos / Precisión)


      temp = 0.5 top-p = 0.95 top-k = 20 min-p = 0.05 repeat-penalty = 1.0

      ------------------------------------------------------------------------------

      QWEN 3.8 27B — PERFIL DE RAZONAMIENTO & CODIFICACIÓN PESADA

      ------------------------------------------------------------------------------

      [qwen3.8-27b] model = /opt/llama- infrastructure/models/Qwen3.8-27B-UD-Q3_K_XL.gguf

      Desactiva "fit" para evitar que capas se carguen en la CPU por un error de

      cálculo automático

      fit = off ctx-size = 73728 context-shift = 1

      MTP nativa del modelo (Decodificación especulativa)

      spec-type = ngram-mod,draft-mtp spec-draft-n-max = 2

      Cuantización de KV (q4_1 nos permite meter contexto de 73k en 16GB de VRAM)

      cache-type-k = q4_1 cache-type-v = q4_1

      Parámetros de presupuesto de pensamiento / razonamiento

      chat-template-kwargs = {"preserve_thinking": true, "reasoning_effort":"medium"} reasoning-budget = 5000

      Batches más chicos para evitar picos de VRAM durante prefills masivos

      batch-size = 1024 ubatch-size = 512

      Ajustes oficiales / recomendados del sampler de cuantización

      temp = 0.65 top-p = 0.95 top-k = 15 min-p = 0.05 ``` submitted by /u/chiribe
      [link] [comments]
      ---|---

    9. 🔗 seanmonstar Tending my little plot of the Internet rss

      As many parts of the Internet continue to get worse, I figured it was time to improve my own little plot.

      I mean, I’m always tinkering, making small tweaks regularly. Did you know I keep /now up-to-date? But anyways, a couple changes here felt big enough to write about.

      My microblog is mine

      I’ve been outputting random microposts since… checks archive 2009, apparently. They were “status updates” back then. With Twitter dying, I started doing that sort of thing on Mastodon, and then on BlueSky. Wherever the people want to be, I suppose.

      At first they were just silly jokes. But eventually, besides announcements, they became ways to express raw (bad) ideas and get feedback. But something about that always bugged me: they were on someone else’s property, and linking to them (let alone finding them again) felt bad.

      So, I own my microblog now. They have their own place on this domain. They get a dedicated RSS feed. And they are included in the main feed (currently prefixed as “Micro” so you know). They get syndicated to those other networks automatically as threads.

      What makes them micro? I don’t constrain myself to just 250 characters or anything. They’re so far about 3 paragraphs. That’s about the size, I aim for, I guess. It let’s me publish thoughts without blocker energy telling me I need to polish it into an essay. It also allows me to output 1 or 2 a week. Or none.

      And I can link to them and build on them. Mine!

      Subscribe via email

      I have improved the subscribe via email option of this site.

      For a long time, “subscribe via email” was easy and nice, provided by Feedburner. When that service was shutdown, I looked for an alternative. Something that was both free and automatically just worked from an RSS feed.

      I’m sorry about that. I picked something horrible, a service I don’t want to provide any further attention. They inject gross click-baity ads inside the emails. I subscribe to myself, and after being repulsed at the last email, I had to fix it.

      I couldn’t find any other service that automatically works from RSS for free. I could pay for a service, but I don’t need to send that much email. And I’m doing this as a convenience, to let users read how they want, not as a business. It’s not a newsletter.1

      So I imported that list to Buttondown. I can copy-paste the markdown of blog posts manually. That’s fine, I don’t write so often to need it to be automated.

      But since it is manual, I can do more. I can also include a list of “recent microblog posts”, now that I own them.

      Anyways, back to continual tinkerage.2

      1. And there’s no way I could subject my readers to Substack or Medium or something. Those sites do not treat readers well. I automatically refuse to read any article on such a site. I assume that if you don’t care about my reading experience, I don’t care enough about your idea. Not sorry. 

      2. Other things I want to improve: a combined blog and micro archive. A tags page. A better chronological story for About. A set of “values” pages. 

    10. 🔗 r/LocalLLaMA Petition to add a rule for people to add their DAMN quant levels to their posts rss

      Every time I see a post about a newly released model, whether it be a comparison or shitting on it, I have to dig through the endless comments to see what quants they used and what their specs were.

      Its quite a common occurrence here in this sub to ask someone that's saying a model is underperforming, and when you ask what quantization they are running they say something like "oh im running q0.1bpw from nobodyknowswhothisguyis".

      Worst offender is with comparison posts. "Comparing the new Qwen3.8-27B to Qwen3.5-9B and the 9B model is better!" I wonder why?

      Sorry for bad england

      submitted by /u/Su1tz
      [link] [comments]

    11. 🔗 r/LocalLLaMA Stripe will reportedly acquire AI gateway startup OpenRouter for $7B+ rss

      another one ..

      submitted by /u/ab2377
      [link] [comments]

    12. 🔗 r/LocalLLaMA …and I’m not afraid of losing my social credits. rss

      …and I’m not afraid of losing my social credits. | submitted by /u/JLeonsarmiento
      [link] [comments]
      ---|---

    13. 🔗 Filip Filmar Bazel all the way down: how I build programmable hardware rss

      This is a description of how I build programmable hardware. Everything that goes into Cocoapuffs, my RISC-V system-on-chip on an Artix-7 FPGA: the RTL, the firmware, the simulations, the synthesis, the bitstream, and the programming of the board, comes out of a single bazel build, from a machine that has nothing installed on it but bazel. The build is hermetic, ephemeral, and reproducible, and it is the same build whether it runs on my laptop, on a virtual machine in the cloud, or in continuous integration.

  4. August 16, 2026
    1. 🔗 Simon Willison Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things rss

      Friday's big release was Qwen 3.8 27B, an Apache 2 licensed 27B parameter vision-capable LLM from Alibaba's Qwen research lab. I've been looking forward to this one: 27B is an excellent size for running a model on a reasonably specced laptop, and its predecessor Qwen 3.6 27B was impressive.

      Qwen's self-reported benchmarks for this model are eye-opening. They show a boost from both Qwen 3.6 27B and the closed-weight Qwen 3.7-Plus, which was one of Qwen's strongest models of any size as recently as May this year. It will be interesting to hear what independent benchmarks have to say about the model.

      I've been running the model on two different machines: my 128GB M5 Max MacBook Pro, and an NVIDIA DGX Spark. On both machines I'm running LM Studio and their 17GB Q4_K_M quantized build. I also tried using llama-server directly on the Spark.

      The default of extra high results in spectacular over-thinking

      Qwen's documentation describes the model as defaulting to xhigh for the reasoning effort, and the LM Studio GGUF I've been trying preserves that default:

      Qwen3.8 comes with official support for reasoning_effort, which can be used to adjust reasoning depth and control cost:

      • xhigh (default): for complex tasks demanding thorough analysis
      • medium: balancing accuracy and speed
      • low: efficient reasoning optimizing for speed and cost

      This is a hilarious default. It's absolutely not a good way to run the model, especially on consumer hardware. I've been finding the results extremely entertaining.

      I quickly ran into problems with LM Studio's default context limit of 8,192 tokens - Qwen was using them all up thinking about even the most mundane of problems. I loaded the model with the full 262,144 maximum context length and that problem went away.

      Here's the pelican riding a bicycle SVG I got from my first attempt with that increased context length. It took 21 minutes to generate, using 22,276 reasoning tokens to produce 3,223 tokens of output. You can read the reasoning trace here.

      A very pleasing image of a pelican riding a bicycle. The bicycle is red and has the correct frame shape. The pelican looks like a pelican and has its wing extended to the handlebars.

      This is by far the best pelican SVG I've been able to generate with a model that runs on a local machine - and this Qwen is pretty small, just a 17GB file on disk. There's a lot to like about this:

      1. The bicycle frame is the right shape
      2. It has legs on each side of the bike - that's very rare
      3. Good, clear pelican pouch
      4. The wings extend to touch the handlebars!
      5. The motion lines are behind, not in front
      6. It has a tasteful background - nice sun, clouds, hill, flowers and grass.

      Was that worth waiting 21 minutes for? Absolutely not.

      Here's that same prompt run with reasoning turned off - transcript here. This one produced 3,715 tokens and took 137s - just over two minutes.

      The bicycle frame is a bad shape now. The pelican remains recognizable but has a less obvious pouch. The feet miss the pedal, there is no attempt to hold the handlebars. And just for completeness, I used OpenRouter to run the same prompt through the much larger Qwen 3.8 2.4T-A95B (released last week) and got this snazzy animated SVG:

      I said Qwen at xhigh has a tendency to over-think things, but how bad really is it?

      I tried a much simpler prompt, again with that default extra high setting:

      draw an svg of a circle

      Qwen's reasoning trace started like this:

      The user is asking for an SVG drawing of a circle. Simple request — but I want it to be a carefully crafted piece. Let me make something that goes beyond just <circle>: a single self-contained SVG file with character — maybe a geometric "circle study," with subtle animation, layered rings, and a distinctive palette.

      Keep the scope right: they asked for an SVG of a circle. So the core is a circle. But I can add craft: concentric guide circles (like a compass/geometry drawing), tick marks, a soft gradient fill on the main circle, restrained ambient motion (a slowly rotating dashed ring, pulsing glow). Respect prefers-reduced-motion? That's more for CSS; SVG SMIL or CSS inside SVG will do.

      Palette options: deep teal ink on warm paper? Or bold vermilion circle on off-white with navy construction lines — Bauhaus/compass-drawing vibe. Let me go with a "geometric study" look: cool slate background, or bright paper white? Paper white is fine if it's not the cream-and-terracotta combo. [...]

      Several minutes later it produced this absolutely beautiful animated circle, which was entirely not what I had asked for!

      My strong recommendation: ignore that default. Run Qwen 3.8 27B on low or even no reasoning levels at first. It's a great model, but wow that default setting is a bad place to start.

      It's very good at bounding boxes

      A fun way to test a vision model is to see how well it can return bounding boxes around items in a photograph. I've seen previous Qwen models deal well with this, so I decided to put it to the test drawing bounding boxes around some pelicans.

      I've seen asking for 0-1000 scale produce good results in the past. I tried this:

      llm -a https://static.inaturalist.org/photos/714731804/large.jpg \
        -m lmstudio/qwen/qwen3.8-27b \
        'Return JSON bounding boxes for the pelicans in this photo, 0-1000 scale for each dimension'

      Here's the reasoning trace, which produced this:

      [
        {"bbox_2d": [195, 290, 370, 780], "label": "pelicans"},
        {"bbox_2d": [445, 320, 675, 850], "label": "pelicans"}
      ]

      This is such a good match. Here are those boxes rendered on top of the photo:

      A photograph of two pelicans on a rocky outcrop, with three other smaller birds. The pelicans both have bounding boxes exactly surrounding them, each with a label that says pelican.

      Building a tool to label bounding boxes

      That visualization of the bounding boxes was taken using a new custom tool that I had Qwen 3.8 27B build for me, running offline on my laptop.

      I forgot to dial down the thinking effort so it was massively over-engineered, but it did manage to produce this full interface from this single prompt:

      [
         {"bbox_2d": [195, 290, 370, 780], "label": "pelicans"},
         {"bbox_2d": [445, 320, 675, 850], "label": "pelicans"}
      ]
      

      Build an HTML page which has an input box for accepting the URL to an image and a textarea for accepting the above style of JSON.

      It appends the image to the page, measures its width and height, then treats the coords in the bbox_2d as scaled from 0-1000 and scales them against the actual width and height, then it renders labelled boxes over the image.

      This screenshot shows one of the features I did not ask for - a demo scene, for if you don't have a photograph to test the tool with:

      Screenshot of bbox·lab, a dark-themed web tool that overlays object-detection bounding boxes on an image, with an input panel on the left and a stage on the right showing two labeled boxes around stylized pelicans in a sunset illustration. Header: bbox·lab — normalized 0–1000 coords → pixel overlay; status indicator: RENDERED · 2 BOXES. Panel 01 INPUT (URL + detections) contains an IMAGE URL field reading data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAA+, a DETECTIONS — JSON textarea reading  {"bbox_2d": 195, 290, 370, 780, "label": "pelicans"}, {"bbox_2d": 445, 320, 675, 850, "label": "pelicans"} , an orange RENDER BOXES button, and dashed boxes labeled DEMO SCENE and CLEAR. Panel 03 STAGE header: display 661 × 661 px · 1 unit = 0.661px x 0.661px · nat 1000×1000. The stage shows a flat-style illustration of two dark pelican silhouettes with orange beaks standing in calm water against an orange-to-purple sunset sky with a pale yellow sun and distant birds; an orange bounding box labeled 1 · pelicans surrounds the left pelican and a cyan bounding box labeled 2 · pelicans surrounds the right pelican. Footer: move the cursor over the image to read grid coords; boxes map 0–1000 → displayed px.

      Here's the relevant segment of the thinking trace, where it decided to draw its own pelicans purely because I had used the label "pelicans" in the example JSON I gave it in the prompt:

      Also a "load sample" that uses a known image? Can't depend on external images, but… the image URL input is user-provided; I could add a "try with sample" button [...] Hmm, I can draw a simple scene on canvas, export it as a data URL, and load it into the image — that's self-contained and demo-able! [...] But the user's coords are for an actual pelican image; a generated placeholder can still demo the scaling. Generate a 1000x1000 placeholder: gradient water + two blob-like "pelican" silhouettes placed at the given bboxes (using the same scale — cute: silhouettes at the exact 0-1000 positions, showing the boxes align). This makes for a fun, self-contained demo. Keep it simple: sky gradient, sun, water, two pelican-ish shapes (ellipse body, circle head, beak). Place at bbox centers.

      (I'm slightly nervous that models around the world might have a bias towards drawing pelicans at any chance they can get, brought on by nearly two years of exposure to my own stupid benchmark.)

      Is all that over-thinking necessary? Maybe it is, at least a bit. I tried with reasoning turned off and got this version, (transcript here), which nearly works but shows the boxes in the wrong place:

      BBox Studio screenshot - a solid UI but the yellow and green boxes do not cover the pelicans.

      So without reasoning it didn't quite one-shot a working tool. I'm sure it could get there with some follow-up prompts, but this is a good example of how reasoning can make a difference.

      Yes, it can drive coding agents

      One of the biggest questions around local models is whether or not they have enough horsepower to successfully run a coding agent loop. Coding agents require long context, strong code generation support and reliable tool-calling. On paper Qwen 3.8 27B has all three of these, so is it up to the task?

      My initial experiments with Pi have been very promising. I chose Pi because it has a shorter system prompt than most other options, making it a better fit for trying out smaller models.

      I configured Pi to use Qwen 3.8 27B running in LM Studio on the Spark (shared via tailscale serve) by adding this to ~/.pi/agent/models.json:

      {
        "providers": {
          "spark": {
            "baseUrl": "https://spark-18b3.tail68a31.ts.net/v1",
            "api": "openai-responses",
            "apiKey": "dummy",
            "models": [
              {
                "id": "qwen3.8-27b",
                "reasoning": true
              }
            ]
          }
        }
      }

      Then ran pi --provider spark --model qwen3.8-27b in my ~/dev/datasette folder and prompted:

      how does auth work?

      After a sequence of reasoning and tool calls that accessed a bunch of different files it produced this reply, which is very solid.

      Just one problem: I wanted to share that transcript. So I pointed Pi and Qwen 3.8 27B at the JSONL transcript file in ~/.pi/agent/sessions/--Users-simon-Dropbox-dev-datasette-- and prompted:

      Write Python code to convert this jsonl to markdown

      And it built and tested this pi_jsonl_to_md.py, which did exactly what I needed. Here's that session transcript, published using the tool that it created.

      The quest for speed

      So far this is all looking very promising. We have a 17GB model that runs on high-end consumer hardware and can write code, drive tools, annotate images and generally do everything that I need from an LLM for getting real work done.

      There's one very significant catch: it feels slow - especially when it starts over-thinking, but even without that it's not particularly sprightly.

      I've been getting around 15-30 tokens a second from LM Studio. That's not terrible, but it's slow enough that it's going to be hard to win me away from hosted API models, which can return results a whole lot faster. Artificial Analysis track token speed and show OpenAI 5.6 Sol at 74 tokens/second and 5.6 Luna at an impressive 184/second.

      The good news is that the community have been exploring ways to speed things up since the model was first released two days ago.

      One of the most promising optimizations is baked into the model itself. Qwen supports Multi-Token Prediction, an architecture trick where a cheaper mechanism guesses several tokens ahead and the main model can then quickly verify if the guesses were correct. This can have quite a dramatic effect on inference performance.

      Based on this tweet from llama.cpp creator Georgi Gerganov I tried running the model with MTP like this on the Spark:

      llama serve \
       -hf  ggml-org/Qwen3.8-27B-GGUF:Q4_K_M \
       -hfd ggml-org/Qwen3.8-27B-GGUF:Q4_0 \
       --spec-default \
       --spec-type draft-mtp \
       --reasoning-preserve

      And sure enough, this gave me a significant boost. I had GPT-5.6 in Codex run a comparative benchmark on the Spark and the --spec-type draft-mtp server outperformed the LM Studio default GGUF by around 72%.

      I expect we'll see a whole lot more innovation around serving this model faster over the next few weeks. The MLX community likely have some tricks brewing as well.

      Some observations

      The fact that a 17GB file can do all of this stuff on my home machines is a miracle. Once again, I'm delighted and amazed at how much progress local models have made this year. A year ago this would have been competitive with the best and most expensive of the proprietary models - today it can run on a capable laptop.

      The only thing holding this back from being a daily driver is performance. It feels pretty slow on both the M5 Mac and the DGX Spark. That's the catch with these dense (non-Mixture-of-Experts) models - they require a whole lot of memory bandwidth to perform well, and neither of the machines I have access to are top performers in that regard.

      The most important thing about Qwen 3.8 27B is what it demonstrates. We can have an open weights general purpose model with a long context, effective tool calling, strong vision ability, and competent code generation, and we can fit the whole thing in just a 17GB file.

      The models at this size continue to get better at an impressive rate. We don't need to spend half a million dollars on datacenter-class hardware just to run a competent model.

      You are only seeing the long-form articles from my blog. Subscribe to /atom/everything/ to get all of my posts, or take a look at my other subscription options.

    2. 🔗 modem-dev/hunk v0.19.0 release

      What's Changed

      hunk-v0.19.0-launch.mp4

      Highlights

      • Install shared extensions from Git, build docked panes and session keyboard modes, and use Hunk's bundled extension-authoring skill by @benvinegar and @mikeclarke in #697, #708, #710, #712, and #717.
      • Guide reviewers to exact code with line navigation and contrast-safe character-range highlights for extensions and live agent sessions by @elucid in #726, #727, and #728.
      • Keep large reviews responsive with in-process untracked-file diffs, active-review syntax caches, and experimental worker highlighting by @benvinegar in #738, #754, and #759.
      • Control the workspace more precisely with configurable files-pane visibility and independent extension panes by @skaragianis and @tridha643 in #648 and #757.
      • Verify release archives with GitHub build-provenance attestations and install Hunk through mise across macOS, Linux, and Windows by @elucid in #714 and #777.

      Full Changelog : v0.18.2...v0.19.0

    3. 🔗 HexRaysSA/plugin-repository commits sync repo: +1 plugin, +1 release rss
      sync repo: +1 plugin, +1 release
      
      ## New plugins
      - [climacros](https://github.com/allthingsida/climacros) (1.0.5)
      
    4. 🔗 r/LocalLLaMA Let’s all thank Georgi Gerganov who gave use llama.cpp rss

      Let’s all thank Georgi Gerganov who gave use llama.cpp | I was looking into the story a bit further earlier. Very interesting. Couldn’t have done it without him submitted by /u/on_line187
      [link] [comments]
      ---|---

    5. 🔗 r/LocalLLaMA Newer commits removed the Qwen 35B rss

      Newer commits removed the Qwen 35B | In this commits, the 35B model was removed. Looks like it's confirming the 35B model won't get released. I think they need to be made aware how big the 35 moe is widely used. Think need to make noise on theyre X, huggingface and online places. If they dont know there's no need to release for people group who dont speak up. submitted by /u/Local-Cardiologist-5
      [link] [comments]
      ---|---

    6. 🔗 Register Spill Joy & Curiosity #95 rss

      In the last two weeks I've shipped: a new experimental provider backend for our orbs, an in-product bug reporting feature (not released yet) including an admin area where we can triage bugs, disk and memory warnings for orbs, visible setup logs when orbs are starting, a full Comet Busters-like game that's hidden as an easter egg on our website, a microphone selector for our dictation features, a new work-in-progress page that explains what orbs are that has a bunch of handwritten text and videos and other stuff I put in there by hand, user preferences for themes, and a few smaller things.

      I also fixed around twenty bugs and removed 5k lines of code that we no longer need.

      "We get it, man, you shippe--"

      Nah, nah, nah! Not the point. The point is this:

      I have not used my local development environment for any of this. I've done all of this remotely, using Amp, in orbs. Everything! Backend for remote machines; messages sent across three services to warn about system resources; landingpage. The freaking game is probably the least surprising thing here, isn't it? And it's a game with custom assets!

      Isn't this wild? No, I know, it is, that's what I'm saying.

      "Surely some things you want to check or test locally, no?" Nah, not really. I mean, yes, that's probably what I would've said half a year ago if you'd told me I won't need my local dev setup anymore.

      Turns out that, no, you don't. You can just ask the agent to give you "irrefutable proof" that something works and if you have an orb and it can do whatever it wants and install whatever it needs it will find a way to give you that proof. Orbs are malleable, the agent can shape them to fit the task by installing and running whatever it needs and and then you get a bespoke made-for-exactly- this-task machine in which an agent can go crazy and if you ask it it will give you a presentation or a narrated video in which it shows by -- frame-by- frame, man! -- that the race condition has been fixed.

      And then, what else do you need your local dev env for? Editing code by hand? Come on, man. Reviewing code deeply? Amp has a diff viewer, so you don't need to do that locally either. And for all of the things I shipped here, I didn't review each line anyway. I do spot checks and make sure the architecture is right, yes, but do I need local tools for that? No. You can ask the agent to help you with reviewing by quizzing you, by giving you diagrams, by showing you a presentation.

      What about the fiddly things? The things you do want to feel your way towards, with your hands? Little bit of padding here, some margin there; now let me flip these two paragraphs and-- ah yes, better. That kind of stuff? That's actually where I'm now experimenting the most because I do have this need to flip words and paragraphs and move stuff around. I want to look at it, change something, look again; undo, redo, change, undo, and back around again.

      But here too the game has changed in a way I find marvelous. Because you can just dictation-dump all your ideas to the agent and hand it screenshots and assets and raw notes and snippets and then ask it to provide you with example pages and 15 different variations of the widget you're interested in, and then you can tweak those and say "this one's good, let's use this one" and you feel like you're the head chef strolling through the kitchen, spoon in hand, tasting the soup over here, tasting the dessert over there, saying "nah" or "mmmmh, good" or "into the trash", and your headless and faceless and bodyless sous-chefs don't mind at all and just do what you say and try again.

      Then weeks go by and you notice you haven't git pulled in a long time and every time you do end up doing that again (due to nostalgia?) maybe use more than one checkout, you notice that it starts to feel… yucky? dirty? unclean?

      Wild times. Exciting times. The models are there now. And if you doubt that, just wait a couple months.

      • New episode of Raising An Agent is out! We recorded this one in-person, in Munich, and talked about everything that was on our mind last week (and this week): orbs, jellyware, why AI by itself doesn't lead to slop, how you need to rethink software now, and, maybe most importantly, the mind-blowing realization of that week in Munich, that no one cares about their local dev env anymore. We all had to wipe our laptops four weeks ago and people said they still haven't ported their dotfiles over and at this point don't care anymore.

      • Speaking of which: I recorded a short video on why orbs aren't "just VMs" and why saying "orbs are just VMs" is missing the mark, just like saying "the cloud is just another person's computer". Also: woo boy, some people are really bothered by product names? Well, too late. It's orbin' time. I get emails from customers telling me they want to get their team "into orbit", others signing off with "happy orbin'!", and customers greeting us in Slack channels with "I love me some orbin' in the mornin'."

      • My teammate Will wrote about how we can push straight to main and still have SOC2. One of the most asked questions we got in the last few months: "Wait, you don't use pull requests? How do you have SOC2 then?" Turns out that SOC2 doesn't require PRs.

      • Very short video in which I show off how I iterate with agents in orbs, working on landingpages and making visual changes, something.

      • Stolen Thoughts - Stealing Reasoning Traces from Proprietary LLM APIs. First of all: wow, what a name, what a website. And then, of course, this is fascinating, isn't it? But I'm not sure whether it's much more than that.

      • Wired also has a write-up on it: A New Trick Reveals AI Models' Inner Thoughts.

      • Read Austin Kleon's Don't Call It Art. Lovely, as expected. If you're in any way interesting in making things or building or writing or just … doing stuff on the Internet: get all of his books. They're very short but very good and very inspiring.

      • It's been a while since I've wanted to access to something this badly: "Cerebras powers GPT-5.6 Sol on Ultrafast mode, delivering up to 750 output tokens per second." Hey, tell the kids to cup their ears real quick. Motherfucking seven hundred and fifty tokens per second. Fucking hell! If there's a sweet angel at OpenAI reading this and can give me access: I will fly to San Francisco and hold your hands and kiss your forehand before I kneel down to thank you and to bless your family and the house they live in and the ground they walk on.

      • Zuckerberg weighs in: "I do not understand why anyone who believes that AI will eliminate most jobs and much of humanity's relevance would rush to build that future." (No, I have not read the whole thing.)

      • The hardest working font in Manhattan. This was long, but soooo good. So good. On a spectrum from "doesn't care about to fonts" on the left to "writes a long and deeply researched article about the history of an unknown font" I'm slightly to the right of center, but I read the whole thing and think you should too if you ever thought "that's a neat font." (Except if that font was Papyrus, of course.)

      • Nail it to the walls: There are no lossless transformations of natural-language text. Very, very, very good. (Sidenote: can you imagine working at a 1000 people org and people use AI to generate Slack messages, emails, PRDs, and slide shows? Yup. Horror stories between two parens.)

      • Craig Mod: A Swarm of Blood Robots. Insert the usual adjectives that I use when talking about Craig Mod's writing: excellent, fantastic, lovely, beautiful. They all apply here too. There are so many things I want to quote here: the section about writing with LLMs, the part about the Weirdness, some of his example projects, the end about the usefulness of these tools. But instead let me just share this one observation: maybe my views on the future of software are so aligned with Craig's (if you go back and read the last twenty issues of this newsletter you'll find that my thoughts on liquid software, jellyware, the future of software, etc. match what he's describing here) because Craig is not part of the software industry and he's not huffing and puffing about how things aren't done properly and he's not stomping his feet about these models being bad at X and Y and he's not stuck in a ten-year old world view of how software's supposed to be built and instead he just has a ton of ideas for things to build and leans into seeing what these models can do and then goes and does it.

      • Sudo Aquarelle, a watercolor simulator. So nice.

      • Finally an end to this stupid argument: "Code was never the hard part" is an insult to all programmers.

      • There is No "Done": Reflections on a Completed AT Thru-Hike. This was great, saying that as someone who's dreamt of walking the AT since he read A Walk in the Woods many, many years ago.

      • I didn't know that the Apple TV has color calibration via iPhone.

      • One of the most beautiful things I've come across this week: Ordinary Abundance. "All the items in this room were once out of reach; some not yet invented, others too rare or costly for the vast majority of people. Today, most of us lucky enough to live with them walk past without a second thought." We'd all probably do well by scrolling through it once a week.

      • Are you a hardcore Rust engineer and want to work remotely with a small and equally hardcore team and do systems- and infrastructure programming? Look no further. I highly recommend working with Nathan and Nick.

      • The Antithesis Principle. This was fascinating. I failed to apply it to every example and got different answers, which makes me think that either (a) the principle is not that clearly defined (possible) or (more likely) that (b) my brain's not wired in this way and I could probably benefit from rewiring it a bit.

      • OpenAI has a friction@ email address employees can use if they feel like they're being blocked.

      • My wife and I were talking about Dolly Parton this week and I said, "Have you ever seen her when she was younger? Or heard her talk?" She said, "No, I haven't." I immediately pulled out my phone and showed her this video.

      • Reminds me: I fell off the wagon again and have watched this video five times in the last 24hrs and I'm about to watch it again, so here, you watch it too. It's only one of the greatest things ever recorded. Danny Carey performing Pneuma.

      • "My dad used to tell me that you could yell at a bear and it would go away. Camping when I was 10, a bear came into the site. Dad got out of the tent and yelled at it, and it just snorted back at him. Dad got back into the tent. 'That's all I got.' A lot of life is like this."

      You should subscribe and tell your friends about this newsletter too, because I know you love it and you keep telling me, but I need the numbers to go up, for the shareholders (and the board):

    7. 🔗 r/LocalLLaMA Qwen3.8-27B vs Qwen3.6-27B writing ray-tracers in BASIC rss

      Qwen3.8-27B vs Qwen3.6-27B writing ray-tracers in BASIC | one of my llm hobbies is re-creating graphics demos i used to write in BASIC in the late 1980s. i slopped together an agentic harness and a basic-to-js transpiler in a web page i've been playing with for a few months. the agent can write basic programs, run them, examine the resulting images, and iterate. qwen3.6 could do a ray-tracer with some user input -- often it got something wrong that it couldn't see/didn't notice, and hence wouldn't fix without further prompting. qwen3.8 typically knocks it out of the park on its own, iterating to a good result. both models are running the unsloth UD-Q8_K_XL quants. i'm pretty happy with 3.8 so far. the user prompt was "write a recursive ray-tracing demo to render three metallic spheres (copper, silver, gold) over a glossy checkerboard plane and under a deep blue sky. use the cook-torrance model to render the spheres." submitted by /u/Ok-Breakfast1878
      [link] [comments]
      ---|---