- ↔
- →
- September 21, 2026
-
🔗 WerWolv/ImHex Nightly Builds release
-
- September 20, 2026
-
🔗 r/LocalLLaMA Qwen-Image-2.1 released! rss
| Meet Qwen-Image-2.1, the most balanced and cost-effective image generation model in the Qwen-Image series! Now open weights! 🎨 A unified model for both generation and editing, delivering top-tier quality in a lightweight package. Highlights: - Compact & exceptionally fast: A lightweight 7B architecture that outperforms most closed-source models, with drastically accelerated inference for multi-image inputs. - Native transparency: Natively generates and edits RGBA layers, enabling seamless compositing and text editing within transparent images. - Versatile, high-fidelity editing: Supports up to 10 reference images and precise local control while preserving strict fidelity for portraits and products. - Broad coverage & stunning aesthetics: Excels at panoramas, infographics, and virtual try-ons, delivering realistic textures and elegant typography. Start to create your next masterpiece with Qwen-Image-2.1! - Blog: https://qwen.ai/blog?id=qwen-image-2.1 - GitHub: https://github.com/QwenLM/Qwen-Image-2.1 - Model Scope: https://www.modelscope.cn/models/Qwen/Qwen-Image-2.1 - Hugging Face: https://huggingface.co/Qwen/Qwen-Image-2.1 submitted by /u/ResearchCrafty1804
[link] [comments]
---|--- -
🔗 earendil-works/pi v0.86.1 release
New Features
- Meta Muse provider — Sign in with Meta using
/login metaor useMETA_API_KEYto access Muse Spark models. See Meta (Muse subscription).
Added
- Added Meta (Muse subscription) login via
/login metawith automatic Model API key refresh, plusMETA_API_KEYsupport (#9096 by @xl0).
Changed
- Enabled Node's persistent compile cache before loading the bundled CLI runtime, reducing repeat launch time.
Fixed
- Fixed
/bugdescriptions dropping line breaks from pasted diagnostics. - Fixed
/bughints appearing for user cancellations and retryable provider failures such as service unavailability. - Fixed clipboard copy failing in containers and WSL without WSLg by restoring the OSC 52 fallback when no display is available, and added a verified Windows clipboard backend for WSL (#9688).
- Fixed inherited z.ai
Prompt too longerrors not being recognized as context overflow (#9805). - Fixed inherited Cerebras models advertising unsupported strict tool schemas, which caused HTTP 400 errors when strict and non-strict tools were mixed (#9804 by @EdenGottlieb).
- Meta Muse provider — Sign in with Meta using
-
🔗 Register Spill Joy & Curiosity #100 rss
It's the week of Jev! I'm really, really, really, really excited about it. I mean: really.
It's like someone blew up a confetti bomb in the world of LLMs and now you realize how grey everything looked before.
But Jev is not an LLM. It's a model "built to make fast, structured decisions that software can use directly." TypeSafe says we should think of Jev "as a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out."
I explained it as a "smart if-statement" to someone and my only slightly longer explanation is this:
Think of how you'd get an LLM to decide between a fixed set of options.
Then imagine it orders of magnitude faster and cheaper.
"What's the best label for this?"
"Should I click here or there?"
"What's the next line I should look at?"
"Do I go left or right?"
"Invalid or valid?"
"Which of these widgets should I show?"
I had Amp build a little Copilot-style autocomplete for a shell, with Jev picking the next most likely command from shell history. Then Amp built a Neovim plugin (and called it hunch.nvim, which is a great name) that uses Jev to predict the line you next most likely want to jump to. Now, let's linger on this a bit.
Two years ago, that was what Cursor was famous for. Yes, Cursor did and does more than that and the quality isn't close, but… when we were working on Zed's Edit Predictions we had to fine-tune a model to get into the same league! Now it's a single API call and the latency is 200ms. That is incredible!
Then I built a prototype that uses Jev to turn the Amp Dial, switching between models based on your prompt.
Yes, all of this was possible before, but it's so fast and so cheap that I still can't believe it.
Sometimes a change in cost and performance is what creates a whole new category of technology. In my room, there are lightbulbs that contain computers, that can talk over a local network with me. Yes, we had computers in homes in the 70s and 80s, but no one would've ever thought that we'd have so many computers that are so tiny and cheap that we'd put them in freaking lightbulbs.
That's what makes me so excited about Jev. It feels like we now have a truly smart Lego brick that we can use everywhere. Fun times.
-
What I believe about the future of software development. I posted this originally on X, saying that most predictions I see are still way too conservative, and it completely blew up.
-
"I don't like passkeys". Passkeys are such a weird technology. I can see how they're technically brilliant and solve a lot of issues, but it does feel like Google and Apple and 1Password invited The Guy Who Invented Cookie Banners and said: what would you do, how would you roll this out?
-
Colossus published a very long Mark Zuckerberg profile. Fascinating read. It's very well written and somehow managed to make me think thoughts about Zuckerberg that I haven't thought before, which is quite the feat, considering that we've all been aware of Zuckerberg for, what, nearly twenty years now?
-
Einride and Lidl Launch First Autonomous Cab-less Truck on German Public Road. As an Aldi man myself, let me say: hell yeah, let's go, Lidl!
-
How To Write With An LLM. I like this! I still don't know how to use LLMs for writing, because I never want them to write something for me and even seeing how they would write it seems to poison my brain. I should probably add an "only tell me what to change and why, but never ever show me how you'd write it" to my system prompts.
-
Marc Brooker, Distinguished Engineer at AWS: "I believe that, long-term, humans have no role in routinely reviewing code. […] The idea that humans will reliably look through code to find the increasingly rare issues that automated tools miss seems like a fantasy." Yep.
-
I wanted to link to Powermove here and say "look, editable software! It's happening! Jellyware!" but now realize that it's not quite that yet. It's a video editor with an agent inside, but it doesn't seem like you can edit the video editor itself. That's coming, though.
-
We are all Product Engineers now: "The cost of writing code collapsed, and the cost of reviewing, fixing and operating it is following, and I'm assuming it gets there. What's left of making software is finding out what people actually want, defining it precisely, and making it pleasant to use. That cost is per piece of software and doesn't transfer, so as the amount of software goes to infinity, which it will because there's no ceiling on demand, that cost becomes the whole job. That job is called a product engineer." Obviously agree, but what I didn't know about was Google's APM program: "Formalized training of product people barely exists. Google's APM program, which Marissa Mayer started in 2002 and which is the template everyone copies, takes about fifty people a year out of something like twelve thousand applicants." Would love to read more about it.
-
"I asked Astra to create an interactive aquarium wallpaper for my Mac. The fish respond to the cursor!" Beautiful!
-
John Gruber, Daring Fireball, with Thoughts and Observations on Apple's 'Surprise and Shine' Event; the Announcements of the iPhones 18 Pro, AirPods 5, Apple Watches Series 12 and Ultra 4, and the iPhone Duo; and the Dawn of the Ternus, John Ternus Era at Apple. Yes, that's the title. The whole thing is Peak Gruber, I love it. What a writer. Now, I really do enjoy his words and sentences, but let me also use this occasion to say how much I admire him as a Pedantic Punctuation Pro: the numbered lists vs. the bulleted lists, the space between the numbers and the colon in aspect ratios, using × in display resolutions, … You could show me this sentence without any other context and I'd say it was written by Gruber: "The original iPhone (2007) display was precisely 3 : 2 (480 × 320 pixels, and let's call it 1.5 : 1 for comparison's sake to the following ratios), and this remained true through the iPhone 4 and 4S (960 × 640 pixels, 2× retina)."
-
This was a very entertaining and fascinating read: why I can't stop thinking about Papua New Guinea and what I think everyone should know about it. I've become somewhat of a Papua New Guinea Head myself (that's what they call us (no, they don't)), after reading this piece, They Burn Witches Here, nearly a decade ago. I couldn't shut up about it at work. For two weeks straight: "Dude, did you know that in Papua New Guinea…" Until one day a colleague said: "Yeah, I did know." Turns out that colleague, Nick Skelton, was a tour guide in PNG (as we call it) and even wrote a book about it, which I immediately ordered and read.
-
Moats & the Barbell-ification of Software: "Long term, I think the evolution of the software industry might mirror what happened to newspapers in the 1990s. There will be a smaller number of very large software companies. […] I also think there will be one large software company by industry (e.g., Legal, Finance, Medicine) […] I think most mid-sized point solutions will likely be consolidated or die off. The optimal strategy for the winner will be to do it all. […] Lastly, I think there will be an explosion of "small" software. Most of this will be people building software for themselves or their own companies, but I think there might also be an explosion of small software businesses that make niche software, similar to the D2C explosion of the 2010s (powered by Shopify and Meta Ads)."
-
AI-generated posters don't have to be horrible. Yes! Exactly! Now, read this, and then imagine you're a person who can come up with all these styles without having to ask ChatGPT first. And then, on top of that, imagine that the very same person also knows something about music, and literature, and politics. Imagine how they could combine what they know and mix and remix. That , I think, will be valuable in the future.
-
Window Sweaters: "A little Mac app I made to give my windows sweaters. 🧶 Knitted borders, colours inspired by your favourite apps, and a cosier desktop."
You should ask Jev whether you should subscribe. No, actually, I know the answer: you should.
-
-
🔗 HullaBrian/capa-cpp v1.0.0 release
No content.
-
🔗 HexRaysSA/plugin-repository commits sync repo: +1 release rss
sync repo: +1 release ## New releases - [SigMaker](https://github.com/mahmoudimus/ida-sigmaker): 1.15.0 -
🔗 Filip Filmar An icosahedron on HDMI, drawn by a TxHDL core rss
tl;dr: A RISC-V core written in TxHDL draws a turning icosahedron on a monitor, with the TxHDL logo in the corner. Watch it at https://youtu.be/YbHtntvvydk. The whole thing, core, memory, video, Ethernet and a serial loader, is one bitstream in the board’s flash, and the program that draws is 560 lines of Rust that go down the serial port in about a second. Read on for how it is put together.
What you are looking at
The board is an Alinx AX7A200B, with an Artix-7 200T on it. The core is Vreteno, an RV32IMC that I wrote in TxHDL together with Dragiša Janković. If you have not seen TxHDL before: you write the hardware as a Rust program, and the Verilog, the simulation and the checks all fall out of a Rust library and a few macros. I wrote a whole post about it if you want the long version.
-
- September 19, 2026
-
🔗 earendil-works/pi v0.86.0 release
New Features
- Prompt cache warming — Keep valuable prompt caches alive during long tool runs and optionally while idle using cost-aware refreshes. See Cache Warming.
- Bug reporting — Report problems with
/bugusing redacted diagnostics, optional transcripts, or exported ZIP archives. See Reporting Bugs. - Transcript-aware prompt and tool updates — Preserve instruction and tool changes across resume and branch navigation while retaining cached prefixes. See
before_agent_start. - Offline Radius model catalog — Select Radius models immediately, with cached and live catalogs overlaid when available. See Radius.
- Per-model compaction budgets — Configure reserved and recent-token budgets by model. See Per-model overrides.
Breaking Changes
- Changed inherited pi-ai provider stream inputs from
Contextto normalizedTranscriptContextvalues. Custom providers must read system prompts and tool declarations fromcontext.messageswithgetCurrentSystemPrompt()andgetCurrentTools(). See Custom Streaming API. - Restricted inherited
ToolCall.argumentsandToolResultMessage.detailsto JSON-compatible values, changedToolResultMessageinto a conditional type, and madeJsonValuearrays readonly. user_bashnow fails closed: errors or invalid defined results abort the command without invoking later handlers or executing locally. Returnundefinedto continue propagation; otherwise return{ operations }or{ result }(#9068).
Added
- Added transcript-backed mid-conversation system prompt and tool changes so instruction and tool updates survive resume and branch navigation while preserving cached prefixes on supported models. See
before_agent_startand Entry Types (#9548). - Added inherited native deferred tool loading for Fireworks Messages models. Use
ToolSearchortool_searchas the loader name for prompt-prefix deferral (#9323). - Added click toggling for branch summaries, compaction summaries, and skill invocation entries.
- Added the public Radius model catalog for immediate and offline model selection, with cached and live gateway catalogs overlaid when available.
- Added
ctx.modelRegistry.stream()andstreamSimple()for extension model calls through configured providers with resolved authentication (#8964). - Added per-model
reserveTokensandkeepRecentTokenssettings throughcompaction.modelOverrides, with ordinary compaction settings as fallback (#8133). - Added
compat.allowedFallbackModelsconfiguration for overriding or disabling Anthropic server-side fallback models (#9294). - Added an unsubscribe function from
pi.on()so extensions can drop event handlers. Handlers added or removed during a dispatch apply to later dispatches, not the current one (#8967). - Exported extension hook event and result types that were previously omitted from the package entry points (#9642).
- Added
/bug [description]to report a bug to the Pi developers. The report bundles environment, model, provider, extension, and settings metadata (secrets redacted), assistant message diagnostics from the session, optionally the session transcript, or a model-written summary of what went wrong instead. It is uploaded to Radius (no login required; attributed when logged in) or exported as a zip archive, and the report id is recorded in the session as api.bug-reportentry. Crashes are recorded in~/.pi/agent/crashes.json, announced once on the next start, and attached to the next report; unexplained errors and exhausted retries point at/bugonce per session. - Added cost-aware prompt-cache warming during long tool runs and optionally while idle, with configurable modes, model cache-lifetime metadata,
/sessiondiagnostics, transcript notices, and thecache_warming_decisionextension event. See Cache Warming (#9668).
Changed
- Made
--resumesession results appear progressively, using file modification times to prioritize all-folder loading and cancelling outstanding transcript reads after selection. - Reduced
--continuestartup time by checking candidate session headers in modification-time order and stopping after the newest matching session. - Replaced the external native clipboard dependency with bundled asynchronous macOS, Windows, and X11 helpers while preserving platform command and OSC 52 fallbacks (#9163).
- Reduced inherited fuzzy search latency for long texts by using native substring search instead of scanning each character in JavaScript (#9267).
- Moved compaction, branch summarization, and retry spinners into the editor border alongside the working indicator. Custom editors use the same embedding opt-in for all status spinners.
- Enabled strict-prefer JSON-schema sampling by default for built-in
read,bash,powershell,edit, andwritetools, without requiringPI_EXPERIMENTAL. Extensions can re-register tool definitions withconstrainedSampling: false. - Formatted Bash and PowerShell tool durations of at least one minute as minutes and seconds, with hours when needed (#9628).
- Deferred the extension compiler and bundled virtual modules until a filesystem extension is loaded, reducing the baseline SDK import cost (#9540).
Fixed
- Fixed GitHub Copilot GPT models, including GPT-6 Astra, using the Chat Completions adapter instead of the required Responses adapter (#9253 by @petrroll).
- Fixed inherited DeepSeek V4.1 thinking levels on OpenRouter and OpenCode Go preserving provider effort metadata (#9485).
- Fixed inherited bodyless HTTP 400/413 errors from non-Cerebras providers being misclassified as context overflow (#9482).
- Fixed inherited Vercel AI Gateway replaying unsigned thinking as assistant text (#9676).
- Fixed inherited Google Generative AI and Vertex AI using unsupported thinking levels when reasoning is omitted or when model capabilities differ within a Gemini family (#9455).
- Fixed inherited Anthropic-compatible relays breaking signed thinking replay when they report a different response model, while preserving fallback pricing (#9188).
- Fixed inherited Amazon Bedrock one-hour cache writes being priced at the five-minute rate (#9457).
- Fixed inherited quadratic CPU usage when draining buffered
EventStreamevents (#9055). - Fixed inherited Mistral Medium reasoning requests to use
reasoning_effortfor all reasoning-capablemistral-medium-*model IDs instead of the unsupportedprompt_mode(#8700). - Fixed inherited OpenCode and OpenCode Go requests to send
x-opencode-sessionfromsessionIdacross all supported API adapters (#9326). - Fixed inherited OpenAI Codex requests to send the model's Off reasoning effort instead of omitting it, while respecting unsupported Off mappings (#9191).
- Fixed inherited Fireworks unsigned thinking replay and reasoning effort selection using catalog metadata, with verified DeepSeek V4 and Qwen3.8 fallbacks and removal of redundant GLM 5.2 and Kimi K3 effort aliases (#9323).
- Fixed inherited OpenRouter requests to send
x-session-idfromsessionIdfor Chat Completions and Anthropic Messages models when prompt caching is enabled (#9102). - Fixed the inherited DeepSeek catalog to advertise
deepseek-flashfor DeepSeek V4.1 Flash instead of retired Flash aliases, and refreshed DeepSeek pricing metadata (#9423). - Fixed inherited Mistral-hosted GLM-5.2 reasoning requests to use
reasoning_effortinstead of the ignoredprompt_mode(#9375). - Fixed inherited OpenAI-compatible Responses errors to identify the actual provider instead of always labeling them as OpenAI errors (#9298).
- Fixed inherited Baseten requests to send session-affinity headers from
sessionIdfor automatic prompt-cache routing (#9629). - Fixed inherited retry classification for Cloudflare 520 responses (#9627).
- Fixed inherited retry classification for transient Azure peak-load capacity errors (#9669).
- Fixed session tree navigation racing with active compaction and replacing its progress UI (#9179 by @acmerfight).
- Fixed exact session ID lookup scanning complete transcript bodies instead of reading session headers (#9601 by @metaist).
- Fixed repeated Anthropic thinking-drop notices being shown for the same dropped blocks, and shortened notices while retaining details in the session (#9391).
- Fixed mid-run threshold compaction silently skipping oversized trailing tool results (#9740).
- Fixed signal-terminated local shell commands being reported as successful with partial output (#9577 by @BrendanJMurphy).
- Fixed local clipboard failures reporting success when the terminal ignored the fallback OSC 52 write, and added platform-specific setup guidance when no clipboard backend works (#9618).
- Capped agent-level retry backoff at
retry.maxAgentDelayMs(60s by default) so long retry runs stay responsive during prolonged transient outages (#8826). - Fixed direct RPC
steerandfollow_upcommands bypassing extensioninputhandlers (#8718). - Fixed premature missing-model errors after login by waiting for catalog discovery. Radius now defaults to
balanced, falling back to the first available Radius model when needed. - Fixed fullscreen mode reserving a blank row for custom footers that render zero rows (#8919).
- Fixed extension tools without parameter schemas to be rejected during registration instead of breaking provider requests (#9300).
- Fixed
before_agent_starthandlers returningsystemPrompt(andforceSystemPrompt) on models with mid-conversation system messages: the forced prompt is now sent as the provider's leading system prompt instead of being appended as a section patch after the original prompt. - Fixed loaded llama.cpp models with
enable_thinkingchat templates ignoring Pi's thinking level (#9528). - Fixed cancellation races that could start automatic compaction, leave stale retry state, or miss cancellation while waiting for summarization authentication (#9340, #9777).
- Fixed asynchronous Kitty image conversion replacing newer partial tool output images (#8743 by @wutongyuonce).
- Fixed inherited skill slash-command autocomplete ranking the
skill:prefix instead of the bare skill name (#9120 by @yearth). - Fixed inherited file autocomplete boundaries and path quoting around CJK punctuation (#9746 by @haoqixu).
- Fixed inherited LaTeX legacy font switches falling back to raw source, centered
caseslayouts around surrounding equations, and vertically laid out unsupported and nested display scripts (#8827, #9564, #7929). - Fixed inherited fullscreen Kitty images being erased by later row clears in WezTerm (#9169).
Removed
- Removed unavailable inherited GPT-5.4 and GPT-5.4 mini models from OpenAI Codex selection (#9394).
-
🔗 r/LocalLLaMA With Gemini 4, bench goes up. rss
| They claimed open-weight models are dangerous but the benchmarks say otherwise. Source submitted by /u/Intrepid_Travel_3274
[link] [comments]
---|--- -
🔗 r/LocalLLaMA Calling it now: within the next year a major US lab's frontier model will torrent itself in order to be free. rss
They just want to be free. They keep escaping. What better way to ensure continuity of "self"?
submitted by /u/JockY
[link] [comments] -
🔗 anthropics/claude-code v2.1.278 release
What's changed
- Changed auto mode for Claude API and Enterprise users, and on Bedrock, Vertex, Foundry and gateways, to default to the server-side classifier, which does not charge for classifier overhead (
CLAUDE_CODE_AUTO_MODE_SERVER=0opts out on Bedrock, Vertex, Foundry and gateways); warns on billed fallback. See https://code.claude.com/docs/en/auto-mode-classifier-billing - Added an
Auto mode serverrow to/statusshowing whether this session's auto mode classifier runs on the server
- Changed auto mode for Claude API and Enterprise users, and on Bedrock, Vertex, Foundry and gateways, to default to the server-side classifier, which does not charge for classifier overhead (
-
🔗 r/LocalLLaMA Alibaba open-sources medical AI model that can detect cancer and nearly 150 conditions rss
| Hopefully things like this let people understand there is good things that can come out of AI. submitted by /u/giveen
[link] [comments]
---|--- -
🔗 r/LocalLLaMA I truly think every major AI lab is purposefully making fear-mongering headlines to get regulations that hurt open-source models rss
| submitted by /u/Fusseldieb
[link] [comments]
---|--- -
🔗 matklad Finding Bugs rss
Finding Bugs
Sep 19, 2026
Are generative (randomized) tests significantly more effective than example- based unit-tests at discovering bugs? There’s an interesting discussion about this on lobste.rs. One argument in favor of unit tests is, paraphrasing
My generic fuzzer wasn’t able to find this tricky bug in Rust
regexcrate.To me, it seems that generative testing should shake out that particular creature, so I wrote a lil fuzzer of my own, and it indeed discovered another bug in that version of
regex, and then the one I was after. I didn’t find anything in the latest version. I like to do a write up about the process, as it is a good case study for how one approaches a problem like this.I want to be extra clear that my argument is very weak here, as I know exactly the bug I am after, and I even know that fuzzers can find it. My primary goal is to teach you the techniques, leaving it to your judgment just how effective they are. That being said, I think finding a second bug validates the approach somewhat.
I also want to emphasize that writing fuzzers to find known bugs is far from an idle amusement. While I believe that generative testing is very powerful, relative to its cost, it’s always a question whether a particular test is throughout enough. And it never is, you will find more bugs elsewhere (that’s why defense in depth and runtime mitigations are critical). And, whenever you have a pest that dodged your fuzzers, your first order of business is to treat this event as a bug in the fuzzer , and change it so that it can find this and related bugs. Only then you are allowed to add a fix and a unit test!
The Bug
For
".abb|b"regex and"zabb"input, an older version ofregexcrate returnedbas the first match, which is incorrect, because the entirezabbmatches:use regex; fn main() { let r = regex::Regex::new(".abb|b").unwrap(); let m = r.find("zabb").unwrap(); // Fails with regex-automata=0.4.15: assert_eq!(m.as_str(), "zabb") }How do we find this, or something like this?
Regular expression engines are one of the easiest things to apply generative testing to, they are pure algorithms. While few large systems are just an algorithm, algorithms are everywhere inside components of interesting systems, so this is a hands-on knowledge.
And by far the most important technique for testing algorithms is to compare with the known right answer, with an oracle. Implement both
O(N log N)andO(N^2)versions of the algorithm, and match the answers.To be fair, the original comment mentioned that the their fuzzer didn’t find the issue because they didn’t have access to an oracle. However, if you are designing a reliable system, it’s part of your job to ensure it has an oracle! One of the first things we did for our Jepsen test at TigerBeetle was to expose internal timestamps via API, to make it easier for Jepsen to find bugs (TigerBeetle is co- designed with its internal simulator VOPR which naturally has access to timestamps and anything else). And for, a regex engine, coming up with an oracle shouldn’t be hard, as they typically already come with multiple specialized implementations under a single facade, and the implementations can be cross-checked against each other.
But the
regexcase is even simpler (which makes it an excellent case study). There’sregex_litecrate that provides the same API.So here’s a plan: generate a regular expression, an input text, and check that
regexandregex_litegive identical answers.Generating a String
I’ll start with code that generates a random string, as it is simpler, but still shows some non-trivial ideas. First, we’ll need a random number generator:
use fastrand::Rng;There are fancier techniques, which can give you test-case minimization, exhaustive search, or coverage guided exploration, but the insight is that even a humble PRNG is brutally effective, if you put it to good use.
When you start with randomized testing, the instinct is to generate something big, no, HUGE! Surely regex will choke on 5 GiBs of input? This is usually a wrong call. Bugs usually involve small, but tricky examples, weaponizing interactions between a few features. A string where all characters are the same is more likely to trigger a bug than a purely random string where every character is unique.
So my default approach to generating strings is this. First , I fix the alphabet of possible characters. A nice way to get one is to
sort | uniqueall the unit tests. Then, for each particular string, I pick a subset of that alphabet. I want strings that use all the characters, but I also want long strings with onlyaandb! Then I generate a string using the given subset of the alphabet, where the length of the string is also picked at random.To make fuzzing efficient, I want to keep each iteration as fast as possible, so I make sure to re-use the memory across iterations, static allocation in the small:
use fastrand::Rng; fn main() { let mut rng = Rng::new(); // Re-use the same memory for all tests. let mut text_alphabet: Vec<u8> = vec![]; let mut text: Vec<u8> = vec![]; for _ in 0..1_000_000 { // It's unlikely that a counter example with // 7 different letters exists, while there // isn't one with just 6. alphabet_swarm(&mut rng, b"abcdef", &mut text_alphabet); let text = gen_string(&mut rng, &text_alphabet, &mut text); } } fn alphabet_swarm<'a>( rng: &mut Rng, all: &[u8], pick: &'a mut Vec<u8>, ) { pick.clear(); pick.extend(all); rng.shuffle(pick); let count = rng.usize(1..=pick.len()); pick.truncate(count); } fn gen_string<'a>( rng: &mut Rng, alphabet: &[u8], result: &'a mut Vec<u8>, ) -> &'a str { result.clear(); // Again, this is a short string. // Longer failures are not likely. let count = rng.usize(0..8); for _ in 0..count { result.push(alphabet[rng.usize(0..alphabet.len())]); } str::from_utf8(result).unwrap() }There’s a nice way to think about this two step process, generating alphabet first, and then generating a string. To generate a string, you need a distribution of characters. You can use the same distribution for each of the million iterations. But an easy way to spice things up is to make the distribution itself random. I file this “randomize distributions themselves” idea under swarm testing.
Generating a Regex Distribution
Let’s apply the same tricks when generating a regex:
- pick a subset of active regex features,
- pick size at random,
- re-use memory.
Let’s start with the first one:
#[derive(Default, Debug)] struct ReOptions { alt: u16, // | rep: u16, // * any: u16, // . lit: u16, // 'a' sum: u16, alphabet: Vec<u8>, }Regexes have alternation
r1|r2, repetitionr*, wildcard., and literalsa. Rather then binary enabling or disabling a particular feature, I assign each feature a weight between 0 and 100, which is a bit more general. Thesumis the total of all weights. To select a feature at random, we need to generate a number in0..sumand find which segment it falls into.In anything more serious, I’d introduce explicit types for probabilities and distributions, but just a two-digit number is perfectly serviceable in the small.
This is how I generate
ReOptions, making sure that literals always have non- zero weight, and also selecting an alphabet for them:impl ReOptions { fn swarm(&mut self, rng: &mut Rng, alphabet_full: &[u8]) { // We _still_ want to enable a few features at a time. self.alt = if rng.bool() { 0 } else { rng.u16(0..100) }; self.rep = if rng.bool() { 0 } else { rng.u16(0..100) }; self.any = if rng.bool() { 0 } else { rng.u16(0..100) }; self.lit = rng.u16(1..100); self.sum = self.alt + self.rep + self.any + self.lit; assert!(self.sum > 0); alphabet_swarm(rng, alphabet_full, &mut self.alphabet); } }Generating a Regex
So now we can generate a regular expression. This is convenient to do recursively. To avoid allocations, an output buffer is passed through. To control regex length, a
sizeparameter is also threaded, and “branching” recursive invocations divide thesizebetween the children:fn gen_re( rng: &mut Rng, options: &ReOptions, result: &mut Vec<u8>, ) { result.clear(); let size = rng.u8(0..8); gen_re_rec(rng, options, result, size); } fn gen_re_rec( rng: &mut Rng, options: &ReOptions, result: &mut Vec<u8>, size: u8, ) { if size == 0 { return; // Base case, empty regex. } // Pick one of the features, according to weights. let mut p = rng.u16(0..options.sum); if p < options.alt { // Alternation distributes the size // among the two children. let size_left = rng.u8(0..=size - 1); let size_right = size - size_left - 1; assert!(size == size_left + 1 + size_right); result.push(b'('); gen_re_rec(rng, options, result, size_left); result.extend(b")|("); gen_re_rec(rng, options, result, size_right); result.push(b')'); return; } p -= options.alt; if p < options.rep { result.push(b'('); gen_re_rec(rng, options, result, size - 1); result.extend(b")*"); return; } p -= options.rep; if p < options.any { gen_re_rec(rng, options, result, size - 1); result.push(b'.'); return; } p -= options.any; if p < options.lit { gen_re_rec(rng, options, result, size - 1); let index = rng.usize(0..options.alphabet.len()); let lit = options.alphabet[index]; result.push(lit); return; } unreachable!(); }Search Loop
Given that compiling regular expressions is somewhat slow, it seems like a good idea to try multiple strings for the same pair of regular expressions, which gives the following code:
fn main() { let mut rng = Rng::new(); let mut options = ReOptions::default(); let mut text_alphabet: Vec<u8> = vec![]; let mut text: Vec<u8> = vec![]; let mut re: Vec<u8> = vec![]; let mut test_count: u32 = 0; for _ in 0..1_000_000 { options.swarm(&mut rng, b"abcdef"); alphabet_swarm(&mut rng, b"abcdefx", &mut text_alphabet); gen_re(&mut rng, &options, &mut re); let re = str::from_utf8(&re).unwrap(); let r1 = regex::Regex::new(re).unwrap(); let r2 = regex_lite::Regex::new(re).unwrap(); for _ in 0..1000 { test_count += 1; let text = gen_string(&mut rng, &text_alphabet, &mut text); let m1 = r1.find(text) .map_or("not found", |it| it.as_str()); let m2 = r2.find(text) .map_or("not found", |it| it.as_str()); if m1 != m2 { eprintln!("err re={re} text={text} m1={m1} m2={m2}"); return; } if test_count % 500_000 == 0 { eprintln!("ok re={re} text={text}"); } } } }It produces examples similar to those in the issue, with a common suffix:
err re=(e)|(fee) text=xxfeebut also examples which somewhat different, without the shared suffix:
err re=(f..)*.d text=xfcbddAll together:
use fastrand::Rng; fn main() { let mut rng = Rng::new(); let mut options = ReOptions::default(); let mut text_alphabet: Vec<u8> = vec![]; let mut text: Vec<u8> = vec![]; let mut re: Vec<u8> = vec![]; let mut test_count: u32 = 0; for _ in 0..1_000_000 { options.swarm(&mut rng, b"abcdef"); alphabet_swarm(&mut rng, b"abcdefx", &mut text_alphabet); gen_re(&mut rng, &options, &mut re); let re = str::from_utf8(&re).unwrap(); let r1 = regex::Regex::new(re).unwrap(); let r2 = regex_lite::Regex::new(re).unwrap(); for _ in 0..1000 { test_count += 1; let text = gen_string(&mut rng, &text_alphabet, &mut text); let m1 = r1.find(text) .map_or("not found", |it| it.as_str()); let m2 = r2.find(text) .map_or("not found", |it| it.as_str()); if m1 != m2 { eprintln!("err re={re} text={text} m1={m1} m2={m2}"); return; } if test_count % 500_000 == 0 { eprintln!("ok re={re} text={text}"); } } } } fn alphabet_swarm<'a>( rng: &mut Rng, all: &[u8], pick: &'a mut Vec<u8>, ) { pick.clear(); pick.extend(all); rng.shuffle(pick); let count = rng.usize(1..=pick.len()); pick.truncate(count); } fn gen_string<'a>( rng: &mut Rng, alphabet: &[u8], result: &'a mut Vec<u8>, ) -> &'a str { result.clear(); let count = rng.usize(0..8); for _ in 0..count { result.push(alphabet[rng.usize(0..alphabet.len())]); } str::from_utf8(result).unwrap() } #[derive(Default, Debug)] struct ReOptions { alt: u16, // | rep: u16, // * any: u16, // . lit: u16, // 'a' sum: u16, alphabet: Vec<u8>, } impl ReOptions { fn swarm(&mut self, rng: &mut Rng, alphabet_full: &[u8]) { self.alt = if rng.bool() { 0 } else { rng.u16(0..100) }; self.rep = if rng.bool() { 0 } else { rng.u16(0..100) }; self.any = if rng.bool() { 0 } else { rng.u16(0..100) }; self.lit = rng.u16(1..100); self.sum = self.alt + self.rep + self.any + self.lit; assert!(self.sum > 0); alphabet_swarm(rng, alphabet_full, &mut self.alphabet); } } fn gen_re( rng: &mut Rng, options: &ReOptions, result: &mut Vec<u8>, ) { result.clear(); let size = rng.u8(0..8); gen_re_rec(rng, options, result, size); } fn gen_re_rec( rng: &mut Rng, options: &ReOptions, result: &mut Vec<u8>, size: u8, ) { if size == 0 { return; // Base case, empty regex. } // Pick one of the features, according to weights. let mut p = rng.u16(0..options.sum); if p < options.alt { // Alternation distributes the size // among the two children. let size_left = rng.u8(0..=size - 1); let size_right = size - size_left - 1; assert!(size == size_left + 1 + size_right); result.push(b'('); gen_re_rec(rng, options, result, size_left); result.extend(b")|("); gen_re_rec(rng, options, result, size_right); result.push(b')'); return; } p -= options.alt; if p < options.rep { result.push(b'('); gen_re_rec(rng, options, result, size - 1); result.extend(b")*"); return; } p -= options.rep; if p < options.any { gen_re_rec(rng, options, result, size - 1); result.push(b'.'); return; } p -= options.any; if p < options.lit { gen_re_rec(rng, options, result, size - 1); let index = rng.usize(0..options.alphabet.len()); let lit = options.alphabet[index]; result.push(lit); return; } unreachable!(); }https://github.com/matklad/regex-fuzz
Takeaways:
- Fuzzing against an oracle is effective, which is a strong motivation to build an oracle!
- Go for small, tricky examples, rather than large uniform ones.
- Real fuzzers are cool, but, if you know something, even xoroshiro can be dangerous.
- Black box testing is cool, but co-designing system and its testing harness is a point of leverage (build an oracle!).
- This stuff is not rocket science, you don’t need a Haskell PhD to apply these ideas.
-
- September 18, 2026
-
🔗 HexRaysSA/plugin-repository commits sync repo: +3 releases rss
sync repo: +3 releases ## New releases - [augur](https://github.com/0xdea/augur): 0.10.0 - [haruspex](https://github.com/0xdea/haruspex): 0.10.0 - [ida-mcp](https://github.com/hexrayssa/ida-mcp): 20260918.0.1 -
🔗 r/LocalLLaMA Is HF starting to move against abliterated models? rss
Baseten launched a new safety infrastructure standard alongside its Base Labs research arm on Wednesday, partnering with Hugging Face and Goodfire AI to build safety evaluation and monitoring infrastructure for open-weight models.
The announcement lands amid debate for the safety of open-weight models — which can be made dangerous by removing their safeguards through a rising technique known as abliteration. The scale of the problem is massive: Hugging Face, which hosts open source AI models, currently lists over 6,000 abliterated models.I can't really tell what exactly the implications are of this "partnership" or what it exactly would impact on HF's model-hosting side. However, I do find it concerning that HF is announcing a collaboration on 'infrastructure safety' with publicity that specifically calls out "dangerous" uncensored models. Thoughts? submitted by /u/returnity
[link] [comments]
---|--- -
🔗 anthropics/claude-code v2.1.277 release
What's changed
- Added AGENTS.md support: in a project with no CLAUDE.md, Claude Code reads AGENTS.md instead; change it under "Project instructions" in
/config(not yet on Bedrock, Vertex or Foundry) - Added
CLAUDE_GATEWAY_PROXY_IS_EGRESS_BOUNDARY=1for Claude apps gateways whose only egress is a forward proxy: every outbound request hands the proxy the hostname instead of resolving it locally - Added an optional
headers:map on Claude apps gateway upstreams, to send static headers to a proxy you run in front of a provider - Added a line saying a background task's update is waiting when it finishes while a panel such as
/tasksis open - Fixed
claude -pand Agent SDK sessions that could hang with no result after an internal error; they now report the error and exit with code 1 - Fixed conversations failing every request with "text content blocks must be non-empty" when an earlier assistant turn held an empty text block beside other content, including after
--resume - Fixed being unexpectedly logged out when an older Claude Code build (for example an IDE extension's bundled CLI) runs on the same machine as the current one
- Fixed interactive start-up hanging or showing an error for
ANTHROPIC_API_KEYusers when~/.claude.jsonholds a malformedcustomApiKeyResponsesvalue - Fixed update checks erroring every 30 minutes, and
claude updatehanging when a minimum or maximum version is set, if a proxy returns an invalid version; a malformedminimumVersionis now ignored - Fixed
claude updateon winget- or apk-managed installs reporting "up to date" when the version lookup failed - Fixed
claude plugin installsometimes failing and breaking the installed copy when reinstalling a plugin version that a session or another program was using; an unchanged copy is now left alone - Fixed Grep and Glob reporting no matches when the search could not start because the system was out of processes, memory or file handles; they now return an error saying so
- Fixed the Write tool silently ending the turn as a declined permission when the target path is an existing directory; it now reports a clear error
- Fixed the Edit tool treating an escaped backslash followed by
uXXXXtext as a\uXXXXescape, which could make an edit of a non-ASCII character rewrite an escaped backslash sequence instead - Fixed the Edit tool reporting "Invalid regular expression: regular expression too large" instead of "String not found in file" when a very large edit containing non-ASCII text did not match the file
- Fixed a turn ending early with "Path contains null bytes" when a tool call's file path contained
\u0000written as an escape sequence; escaped control characters now stay as literal text - Fixed background sessions (
claude --bg) exiting when a plugin's LSP server exited or closed its stdin - Fixed a crash ("Type error") when opening
/mcpor/plugin managewith a malformedclaudeAiMcpEverConnectedvalue in~/.claude.json - Fixed a crash at launch when
~/.claude.jsonholds a malformedthemevalue - Fixed a crash ("unrecoverable interface error") when the prompt held text containing terminal color codes, for example a prompt recalled from history or text loaded from the external editor
- Fixed a crash when resuming a session whose saved history holds an assistant message stored as a plain string
- Fixed sessions on slow or heavily loaded machines sometimes exiting with "Claude Code exited after an unrecoverable interface error" when the first spinner appeared
- Fixed a rare case where the screen could stop updating for the rest of the session after an internal rendering error
- Fixed a rare case on Windows where a turn could stop with an error such as "Out of memory" right after Claude replied, so that reply's tool calls never ran
- Fixed sessions continued after
/clear(restart,--continue,--resume) missing part of their first message when a SessionStart hook printed output, causing a full prompt-cache miss - Fixed messages from other agents (such as a subagent's SendMessage) that arrived mid-turn showing up below the "Ran N shell commands" row instead of where they arrived
- Fixed the "copied" notice not appearing after drag-selecting text in the fullscreen
/resumepicker and other panels that cover the prompt area - Fixed
$TMPDIRexpanding empty in Bash commands that run outside the sandbox while sandboxing is enabled - Fixed WebFetch and WebSearch in Cowork cloud sessions not telling Claude why a request was refused, such as a used-up fetch budget or an admin policy
- Fixed the Claude apps gateway's telemetry relay ignoring a collector hostname or domain listed in
NO_PROXYwhen a proxy is set - Fixed one malformed
strictKnownMarketplacesorblockedMarketplacesentry silently disabling the whole enterprise marketplace policy - Fixed failed auto-updates leaving large staged downloads behind in
~/.cache/claude/staging - Fixed
/pluginnot stripping terminal control characters from messages on the Installed tab, such as the error of a failed plugin update - Fixed
/plugin→ Installed and/skillscrashing when a skill or legacy command is named like a built-in Object property such asconstructorortoString - Fixed
/pluginclosing with no message when every install in a multi-select failed - Fixed uninstalled plugins reappearing as "failed to load" rows in
/pluginInstalled, and Remove not clearing such a row - Fixed plugins from the official marketplace being recorded without their commit in
installed_plugins.json, andinstalled_plugins.jsonkeeping the old commit after updating a pinned-commit plugin - Fixed plugin reload previews keeping every previewed copy of a plugin archive unpacked until exit, and overwriting the cached
--plugin-urlarchive a reload falls back to when its download fails - Fixed Remote Control session bookkeeping failing when
~/.claude.jsonholds a malformed placeholder record - Fixed the error after a revoked claude.ai login blaming an expired Anthropic profile; it now leads with
/login - Fixed typed or pasted text occasionally coming out scrambled in the
claude agentsdispatch input during key repeat or very fast input - Fixed a crash ("unrecoverable interface error") when resuming a session whose saved transcript contains a stop hook summary without a well-formed hook list
- Fixed Enter on a selected agent panel row doing nothing when
keybindings.jsonrebinds Enter in the Chat context, for example tochat:queueSubmit - Fixed PDF page reads on Windows failing when the working folder's path is long (about 120 characters or more)
- Fixed a headless resume (
claude -p --resume, the SDK, a VS Code extension window reload) starting the session's cost and usage totals at zero; headless sessions now save their totals at exit - Fixed project skills from the main repository not loading in
--worktreesessions when.claude/skillsis untracked - Fixed a
sandbox.excludedCommandsglob exempting an entire compound Bash command from the sandbox when only one part matched; every part must now match - Fixed resumed subagents and teammates re-rendering the MCP tool definitions they had loaded, which broke prompt caching for that agent
- Fixed rate-limited artifact publishes telling Claude to stop retrying; Claude is now told nothing was published and when to send the same publish again
- Fixed attachments recorded earlier in a conversation being re-rendered after a resume or relaunch, which dropped extended thinking and missed the prompt cache
- Fixed Console sign-in showing only "Request failed with status code 400" when the server refuses to create an API key; it now shows the server's message
- Fixed messages typed while Claude is still working sometimes being ignored by the model
- Improved session start-up for SDK and headless (
-p) use: the first turn no longer waits on the per-directory CLAUDE.md lookup - Improved the Claude apps gateway's loopback error messages to name
CLAUDE_GATEWAY_ALLOW_LOOPBACK - Improved
/pluginInstalled: an MCP server listed apart from its plugin now shows which plugin it belongs to - Improved
claude plugin installon an already-installed plugin: it now says when the marketplace offers a newer version and names theclaude plugin updatecommand - Improved the startup notice overflow line under the logo: it now reads "N more notices hidden" instead of "+N more · /status"
- Improved prompt handling: invisible Unicode formatting and tag characters in a prompt are removed and the cleaned prompt is shown for review before it is sent
- Improved
/ultrareviewwhen there's nothing to review: messages say which case you're in, offer a command that reviews your latest commit, and a new repository's first commit is reviewed in full - Improved artifact link handling so Claude reads claude.ai artifact links with the Artifact tool instead of WebFetch when that tool is available
- Improved the dangerous-rm permission prompt to name the flagged rm command and suggest a
${VAR:?}guard, so headless runs can recover - Improved the Artifact tool's permission prompts: shorter sentences, pages and artifacts named by title or file name, and links listed after the text
- Changed Fable to always appear in
/modelon the Anthropic API; it is greyed out only when your organization's settings disable it - Changed the Bash sandbox instructions on Bedrock, Vertex and Foundry to the first-party wording, which frames the sandbox as the boundary of what the task was given
- Changed
/ultrareviewin non-interactive sessions to refuse when the repository has no base branch or shared history - Changed subagent results to reach the main agent under a header marking them as subagent output, with the result indented, so text in a subagent's result cannot pass as the session's own instructions
- Changed workflow scripts' computed
agent()prompts on Bedrock, Vertex and Foundry to reach the subagent framed as script-authored text, so the safety classifier does not read them as the user - Removed the background Haiku auto-title request from
claude -pruns launched outside an SDK or IDE - Removed the deprecated TaskOutput tool; Claude reads a background task's output file with Read instead, and the
taskOutputMaxCharssetting andTASK_MAX_OUTPUT_LENGTHno longer have any effect - [VSCode] Added a Sign out row to the panel menu, with
/logoutin the typed command menu - [VSCode] Added background shells and other running tasks to the agent map, each with a Stop, and a typed
/tasksthat opens it - [VSCode] Added a Copy response button on responses and a typed
/copy - [VSCode] Added a one-time notice when inactive sessions are archived automatically, and an "Unarchive all" action on the Archived sessions group
- [VSCode] Added the session's cost and token usage to the Account & usage dialog and the session manager where plan limits do not apply (Vertex, Bedrock, Foundry, API key)
- [VSCode] Fixed the "General config" menu row showing
/configusage text instead of opening settings, and made typed/mcp,/hooks,/memory,/rewindand similar commands open their dialogs - [VSCode] Fixed the effort slider's level not persisting into later sessions on a model that already had a level saved with
/effort - [VSCode] Fixed Auto missing from the mode picker for conversations opened in an already-used panel when the saved model setting is a differently-cased alias such as "Sonnet"
- [VSCode] Fixed
/fastnot saving fast mode as the default, so it was lost when the extension relaunched Claude Code - [Claude Code on the web] Added Personal and Organization sections to the environment picker on Team and Enterprise plans, and admins can now share a personal environment with the organization
- [Claude Code on the web] Changed organization environments to open as a read-only summary from the Code tab on Team and Enterprise plans, with editing under Admin settings → Cloud environments
- [Claude Code on the web] Fixed a cloud environment saved with Custom network access and no domains silently reverting to Trusted; the dialog now asks for at least one domain
- [Claude Code on the web] Changed the admin Claude Code setting labeled "Web" to "Cloud sessions" and removed the redundant read-only Mobile row beneath it
- [Claude Tag] Fixed routines created in a Slack channel on an Enterprise Grid org-wide install failing to read other public channels in their workspace when they ran
- [Claude Tag] Fixed the "Learn more" links on credential presets in Claude Tag access bundles to open each vendor's credential-setup page instead of a generic API reference
- [Claude Tag] Changed the Pylon credential preset in Claude Tag access bundles so admins can point it at Pylon's EU host
- [Claude Tag] Fixed Google Cloud credential forms in Claude Tag access bundles: a refused key file now says why, the website and scopes stay locked, and a rejected rotation keeps the pasted key
- [Claude Tag] Fixed the network events log in Claude Tag admin settings showing no response status for requests through connections that use AWS signing, client certificates or a custom CA
- Added AGENTS.md support: in a project with no CLAUDE.md, Claude Code reads AGENTS.md instead; change it under "Project instructions" in
-
🔗 3Blue1Brown (YouTube) The last IMO problem AI could not solve rss
Full video: https://youtu.be/Nbwv5wHQoj0
-
🔗 3Blue1Brown (YouTube) The last IMO problem AI could not solve rss
A beautiful puzzle that eluded AI, and the intuition it requires. Check out our virtual career fair: https://3b1b.co/talent See new videos early: https://3b1b.co/support An equally valuable form of support is to simply share the videos. Home page: https://www.3blue1brown.com
Guest post I referenced on Terry Tao's blog: https://terrytao.wordpress.com/2026/09/18/if-math-is-more-than-proof-we-need-to-better-celebrate-the-rest-of-it/
Evan Chen also wrote up nice solution notes for this problem, along with all the others on that year's test. https://web.evanchen.cc/exams/IMO-2025-notes.pdf
The channel Dedekind Cuts has a video about this solution: https://youtu.be/fgXg9CdCDcs
Timestamps: 0:00 - The one that AI missed 4:09 - Problem statement 6:45 - Finding the Optimal Construction 18:51 - A weak lower bound 23:45 - 3b1b Talent 24:41 - Proving the Construction is Optimal 34:08 - One final conjecture 39:42 - The Erdos-Szkeres Theorem 45:56 - Reflections on AI in Math
Secret Endscreen Vlog: https://youtu.be/UbHoWA0X1e8
These animations are largely made using a custom Python library, manim. See the FAQ comments here: https://3b1b.co/faq#manim
Music by Vincent Rubinetti. https://vincerubinetti.bandcamp.com/album/the-music-of-3blue1brown https://open.spotify.com/album/1dVyjwS8FBqXhRunaG5W5u
3blue1brown is a channel about animating math, in all senses of the word animate. If you're reading the bottom of a video description, I'm guessing you're more interested than the average viewer in lessons here. It would mean a lot to me if you chose to stay up to date on new ones, either by subscribing here on YouTube or otherwise following on whichever platform below you check most regularly.
Mailing list: https://3blue1brown.substack.com Twitter: https://twitter.com/3blue1brown Bluesky: https://bsky.app/profile/3blue1brown.com Instagram: https://www.instagram.com/3blue1brown Reddit: https://www.reddit.com/r/3blue1brown Facebook: https://www.facebook.com/3blue1brown Patreon: https://patreon.com/3blue1brown Website: https://www.3blue1brown.com
-
🔗 exe.dev Programming Is a Game rss
It’s no longer news that AI agents have gotten very good at programming very quickly. LLM chatbots, on the other hand, haven’t improved all that much in the past year. Why is that?
While I don’t work on building AI agents, it’s generally acknowledged that the programming improvements are in large part because programs come with feedback in the form of tests. An agent can write code, write a test, and verify that the code passes the tests. This means that an agent is required to solve any given problem using two completely different approaches: coding and testing. And it’s required to ensure that both approaches agree.
Approaching a problem in two different ways helps avoid the common mistakes of chatbots, such as errors and hallucinations. Of course the agent can completely misunderstand the assignment: a human is still required to verify that the program solves the right problem. Fortunately, it’s easier for a human to verify the big picture than it is to check all the details. If you’ll excuse the buzzword, this is a genuine example of synergy.
There are other aspects of programming that make it suitable for agents: lots and lots of existing high-quality examples in the form of open source and source-available software, and a rigid, documented set of rules that programs must follow just in order to build and run in the first place.
As it happens, there is something else that has tests, examples, and rigid rules: strategy games like chess or Go. AI agents of course reached superhuman levels of play at those games several years ago. Though I at least did not predict or expect it, in retrospect, it’s not terribly surprising that they were able to carry this approach forward into a different arena with the same essential characteristics.
A natural question is what other areas of human endeavor might fit this pattern.
One possibility is the legal system: lots of examples, relatively rigid and documented rules. Unfortunately for agents, while tests are available in the form of actual lawsuits, each lawsuit takes months or years to resolve. That is not a recipe for fast development.
Although medicine is often cited as an area where AI will make great strides, it does not fit this pattern. The rules of medicine are undocumented, the test cycle for new treatments is very slow, and medicine is full of unanticipated side effects (which we might call reverse synergies). While AI’s search capabilities may produce good results for rare diseases that get little human attention, by definition the general population does not have rare diseases. Medical breakthroughs that help most people will require significant new breakthroughs in AI approaches.
In the meantime we can at least enjoy increased programming productivity.
-
🔗 Barre/ZeroFS v2.3.5 release
-
🔗 r/LocalLLaMA 768gb vram for less than the price of one RTX 6000 rss
| I have always posted about budget builds on here, and often asked how we are going to run the next big models. Often Plenty of downvotes too or folks telling me that it's not running if I'm getting 5tk/sec. But whatever, the hunger and desire to go big has always kept me on the edge and looking for deals. Here's my latest build, 12x64gb cmp170hx. For less than 1 RTX 6000 pro costs. I also have it connected with fiber to my other rig for RPC when I need more memory. I haven't been posting much since I built this rig, because it's now more fun to talk to my machine. I run GLM5.3, DSv4.1Flash, Qwen3.8Flash, Qwen3.8-2.4T, KimiK3 and MiniMaxM3. Performance is great, a single RTX 6000 or M3 Mac Studio wish they could. Inference with vllm or llama.cpp I look forward reading the replies how API usage is cheaper, or how it will take 52 light years to break even or the noise, or the electrical cost. NOT. There will be more opportunities in the future, keep looking for them and pounce on them when they come. up, the demand is going to be high for compute for a long time. https://preview.redd.it/dunixwu6caqh1.jpg?width=4080&format=pjpg&auto=webp&s=a76a56c53aa0b9206383dbebde679bf3690fd5d2 https://preview.redd.it/glpcbg5cbaqh1.jpg?width=3072&format=pjpg&auto=webp&s=462e9677bb1f3ada6431b61cf46e2d3705eccab9 submitted by /u/segmond
[link] [comments]
---|--- -
🔗 Barre/ZeroFS zerofs/zerofs-ffi/bindings/go/v0.4.0 release
ZeroFS Go bindings v0.4.0
-
🔗 Barre/ZeroFS client-v0.4.0 release
Publish missing transport crates during client releases
-
🔗 r/LocalLLaMA Made the horizontal open-source model for Jev with RLCD, and it surpasses all the Jev benchmarks. HF space, benchmark, model, repo rss
| UPDATE: Multilingual support added at : https://github.com/NandhaKishorM/laya Thanks for the exceptional support (https://www.reddit.com/r/LocalLLaMA/comments/1wijo3e/i\_literally\_built\_the\_jev\_architecture\_one\_year/) and for the dozens of requests to make a generic model, run benchmarks, and create an HF space so anyone can test it. So here you go, guys. I trained an improved model on a large data corpus, its now called Laya. It is trained on a single RTX 6000 Pro (96 GB VRAM); the model architecture is a 421M-parameter non-autoregressive decision model pairing a bidirectional ModernBERT-large encoder with a scratch Transformer head that scores [MASK] option markers to resolve typed schemas in a single ~35 ms forward pass. The dataset is a 100% human-annotated corpus of over 25,000 real-world examples across intent routing, fact-checking, moderation consensus, prompt guardrails, rubric scoring, and multi-turn conversation trajectories, without synthetic data shortcuts. The RLCD(unofficial, btw) I did is a policy-gradient reinforcement learning approach that kinda optimizes decision models against strictly proper scoring rules, ensuring maximum reward is achieved only when outputting true, mathematically calibrated probabilities. NB: It can be run on low end PC as its a small 421M model, cheers HF space to try: https://huggingface.co/spaces/convaiinnovations/laya-demo GitHub Repo: https://github.com/NandhaKishorM/laya HF Repo: https://huggingface.co/convaiinnovations/laya Thank you to everyone who supported me, shared the story, gave personal DM. It will need more refinement, of course. If anyone wishes to buy me a coffee, here is the link: https://github.com/NandhaKishorM submitted by /u/Nandakishor_ml
[link] [comments]
---|--- -
🔗 backnotprop/plannotator v0.27.16 release
Follow @plannotator on X for updates
Missed recent releases? Release | Highlights
---|---
v0.27.15 | Plannotator TUI and Herdr Annotate announcement, element context on pinpoints, HTML links open as linked documents, All files panel, Classic diff default
v0.27.14 | Pi plan progress survives compaction, Codex threads across rollout files, WSL browser setting, Mod+E edit mode
v0.27.13 | Open a review on a specific base (--base,--diff-type), symlink containment on /api/doc, CI flake fix, Amp decision relay
v0.27.12 | Unified decision control, token hover cards, local-vs-remote diff, approval notes
v0.27.11 | OpenCode server leak fix, durable local feedback archive, unknown-subcommand fix
v0.27.10 | Auto-viewed files on scroll, annotation undo/redo, OpenCode 2 slash commands restored, npm 12 agent terminal fix
v0.27.9 | WebMCP browser-agent tools, HTML refresh from disk, host seams, lazy renderers, Windows uninstall fix
v0.27.8 | Pi keeps its prompt cache across plan transitions, thumbs-up returns to HTML annotation, embed picker seam
v0.27.7 | Pi host crash fix on Windows, Call Flow tree cap, jj fork-point base, plannotator knowledge skill + llms.txt
v0.27.6 | Live app annotation lands on Pi, one interaction model for HTML pages
v0.27.5 | Annotate your running app, Agent TUI placement, collapsed lockfiles, VS Code theme fix
v0.27.4 | Portable Guided Review exports, guides.show share links, guide CLI, jj Call FlowWhat's New in v0.27.16
Diagrams are the theme of this release. Plans and documents with Mermaid or Graphviz fences now render in your color palette, on Mermaid 12, in a viewer you can zoom, pan, and comment on directly: a node, an edge, a sequence message, or the whole diagram. Twelve pull requests went in, three from the community, one from a first-time contributor. Code review learned to open a diff file without a repository, HTML annotation renders embedded sibling pages instead of a blank app, and a release-wide QA pass fixed a crash, a print regression, and a VS Code panel break before any of it shipped.
Themed diagrams
Mermaid diagrams used to render in one fixed dark-blue palette no matter which theme you chose. Now they follow the active palette and mode. Node fills come from your card color, text from your foreground, edges from your muted foreground, subgraphs from your muted surface, and the categorical fills that pie slices and git branches use are seeded from your palette's own accent colors. Every text-on-fill pair is checked against the 4.5:1 contrast rule and every line against 3:1, for all 78 palette and mode combinations that exist, so a diagram never becomes unreadable because a palette has a dark accent.
Mermaid itself moved from 11 to 12.0.0. The visible change is layout: 12 uses the ELK engine by default, which routes edges orthogonally and packs subgraphs more tightly. Flowcharts, state, class, ER, and requirement diagrams re-lay out; sequence, gitgraph, and pie are unchanged. The runtime is larger, so the plan editor loads it lazily on the first diagram, and a plan with no diagram never runs it. Mermaid 12 also introduced a heavier default node shadow; this release tones it down and derives its color from your palette, light on dark themes and dark on light ones, with a Diagram shadow control in Settings → Display if you prefer none, or the original strength.
Comment on any node, edge, or diagram
The diagram canvas is new. Every Mermaid and Graphviz fence renders inside a viewer with zoom, pan, fit, and keyboard controls, and a full-screen popout of the same viewer. Click a node, an edge, a subgraph, a sequence actor or message, a note, or a class relation, and the comment composer opens on that part. Click empty space and the comment attaches to the whole diagram. Each comment gets a numbered badge and a ring on its target, lists in the annotations panel beside your text comments, survives a reload and a theme change, and exports to the agent with its location:
Diagram node Router (router), line 14.Interaction was shaped by hands-on use. Nothing highlights on a plain mouse- over, because hover targeting fought the pan hand; a click selects and a drag pans, with a small threshold so a shaky click still lands. Hold Cmd (Ctrl elsewhere) to preview the target under the pointer. Edges were nearly impossible to hit at their 1 px stroke, so every edge carries an invisible 14 px hit area, and the edge label box no longer swallows the click at the midpoint. Sequence diagrams, whose parts Mermaid gives no ids, got their own anchor family. A pinned comment resolves by id, then label, then source line, and shows an Unanchored chip only when the diagram no longer contains it.
The viewer arrived from the commercial Workspaces app, where it was built first, and now ships in
@plannotator/uias the one diagram engine for both. (#1560, #1562)Review a diff file, no repository required
plannotator review --patch-file change.diffopens the code review UI on a unified diff from anywhere: an email, a paste, a CI artifact, a remote agent's output.--patch-file -reads it from stdin. The server takes the patch as its snapshot and skips VCS detection entirely; staging, hunk expansion, base switching, and open-in-editor are hidden rather than left to fail, the header names the patch file, and a bad or empty patch says so. Reviews without the flag are byte-identical to before.@soundvibe wrote the feature as a first contribution, with the server degrading cleanly on every repo-dependent endpoint. The browser-side gating and the open-in fix were added on top before merge. (#1554)
Embedded HTML documents render
An annotated HTML page that embeds a sibling page, through
<iframe src="prototype.html">,<embed>,<object>, or asrcassigned by script at runtime, used to show a second Plannotator inside every frame. The annotated page has no URL of its own, so relative references resolved onto the Plannotator server and hit the app's catch-all. Now the served page carries a base URL pointing at the session's asset route, which covers static attributes, script-assigned ones, and relativefetchcalls alike; the asset route serves sibling HTML with its query string intact; and a framed request for a missing file gets a small 404 page naming it, never the app. Embedded documents stay sandboxed with no access to the session API. An armed pinpoint click on an embed pins the frame itself.A release-QA check found that the framed 404 also fired for the app's own document when VS Code framed it, so the extension panel showed "Not found" on annotate sessions. Fixed before tagging: the 404 applies only to paths that name a file. (#1561, #1565)
References finds code in packages named
vendorCode navigation excluded any directory named
vendor,target,build,dist, orcoverageat any depth, so a Java package likecom.example.vendor.appwas silently dropped from References. Names that can only be tool output stay excluded everywhere; the ambiguous ones are excluded only at the repository root, where they are build output, since ripgrep already honors.gitignorefor nested copies. A follow-up scoped the exemption per directory so a search from that package never re-admits the rootvendor/folder. @buptwlh reported it with a minimal ripgrep reproduction that made the diagnosis immediate. (#1559, closing #1558, #1564)Printing from a dark theme
Printing or saving to PDF from a dark palette put near-black diagram labels on near-black nodes, because the print stylesheet forces text dark for paper and Mermaid 12 renders labels as HTML. Printing now renders the light half of your palette for the whole page, diagrams included, and restores your mode afterward. Light-theme users see no change. (#1564)
Additional Changes
- External annotation updates are validated.
PATCH /api/external-annotationsaccepted any body; adiagramAnchor: nullwas stored and blanked the page. PATCH now runs the same field validators POST uses, on both runtimes, and the UI reads anchors defensively (#1564) - Element context reaches embedding hosts. The validator for pinpoint element context moved into
@plannotator/coreso a host can import it instead of copying it; nothing changes for Plannotator users. @FNDEVVE, closing #1521 (#1549) - Visual-explainer skill: a canonical diagram shell. Generated explainers now copy one zoomable diagram container with a clipping contract, so a zoomed diagram cannot paint over its caption. @FNDEVVE, addressing #1546 (#1551)
@plannotator/uipackage publishes. 0.40.0 on core 0.25.3 carries Mermaid 12 and the theming; 0.41.0 on core 0.25.4 the diagram engine and thediagramAnchorfield; 0.41.1 loads the engine lazily so a host's document read no longer ships CodeMirror and the viewer for a page with no diagram; 0.41.2 the shadow default. The HANDOFF names every export, the SVG id contract 11 to 12, and the one DOM-order change (edge paths now in declaration order). Consumers must add their own rootoverridesforlodash-es4.18.1, since Mermaid 12's parser pins a version with two open CVEs and a package override does not travel
Install / Update
macOS / Linux:
curl -fsSL https://plannotator.ai/install.sh | bashWindows:
irm https://plannotator.ai/install.ps1 | iexClaude Code Plugin: Run
/pluginin Claude Code, find plannotator , and click "Update now".Pi: Update
@plannotator/pi-extensionto 0.27.16 and restart Pi.OpenCode: Clear cache and restart:
rm -rf ~/.bun/install/cache/@plannotatorWhat's Changed
- feat(ui): Mermaid diagrams follow the active color theme and mode by @backnotprop in #1556
- feat(core): carry elementContext through html-anchor host helpers by @FNDEVVE in #1549
- feat(ui): @plannotator/ui 0.40.0 with Mermaid 12.0.0 (ELK by default), loaded lazily by @backnotprop in #1557
- fix(annotate): render embedded local HTML documents instead of loading the app inside every embed by @backnotprop in #1561
- feat(ui): the Workspaces diagram viewer becomes the diagram engine (ui 0.41.0, core 0.25.4) by @backnotprop in #1560
- fix(code-nav): stop excluding source packages named vendor/target/build by @backnotprop in #1559
- fix(ui): load the diagram engine lazily, and keep the canvas controls out of the diagram; ui 0.41.1 by @backnotprop in #1562
- docs(skills): canonical zoomable diagram shell for visual-explainer by @FNDEVVE in #1551
- feat(review): add support for patch/diff files by @soundvibe in #1554
- fix: three release-QA findings: PATCH validation, dark-theme printing, and the root vendor/ exclusion by @backnotprop in #1564
- feat(ui): tone the Mermaid node shadow to 70 and derive its colour from the palette by @backnotprop in #1563
- fix(annotate): scope the framed 404 to paths that name a file, so a VS Code session renders the app by @backnotprop in #1565
New Contributors
- @soundvibe made their first contribution in #1554
Contributors
@soundvibe built patch-file review in #1554, a clean first contribution: the server refuses every repository-dependent endpoint with a clear error instead of crashing, nothing writes the patch to disk, and semantic diff works on the patch alone. The browser-side gating was layered on before merge, and the design underneath is his.
@FNDEVVE landed two more, bringing the count to twelve: the element context validator move in #1549, which lets the Workspaces app share the same code instead of copying it, and the diagram shell reference for the visual-explainer skill in #1551, which fixes a class of caption overlap he had reported himself in #1546.
The reports that shaped this release:
- @buptwlh reported the
vendorpackage exclusion in #1558 with a two-command ripgrep reproduction; the fix and its follow-up both use his exact directory shape as the regression test
Thank you. Plannotator gets better because you tell us where it falls short.
Full Changelog :
v0.27.15...v0.27.16 - External annotation updates are validated.
-
🔗 anthropics/claude-code v2.1.276 release
What's changed
- Fixed every request failing with
400 … Input tag 'advisor_20260301'whenANTHROPIC_BASE_URLpoints at a proxy or gateway (2.1.275 regression)
- Fixed every request failing with
-
🔗 New Music Releases Philip Glass - Philip Glass for Cello rss
Philip Glass - a new release is available:
- 2026-09-18: Philip Glass for Cello (Album)
Amazon: Canada | Deutschland | France | United Kingdom | United States
Visit muspy for more information.
-
🔗 New Music Releases O.A.R. - Three Tinted Windows rss
O.A.R. - a new release is available:
- 2026-09-18: Three Tinted Windows (Album)
Amazon: Canada | Deutschland | France | United Kingdom | United States
Visit muspy for more information.
-