- ↔
- →
- October 09, 2026
-
🔗 HexRaysSA/plugin-repository commits sync repo: +4 releases, -2 releases rss
sync repo: +4 releases, -2 releases ## New releases - [ida-bochs-binaries](https://github.com/hexrayssa/ida-bochs-binaries): 2.0.0, 1.0.5 - [ida-mcp](https://github.com/hexrayssa/ida-mcp): 20261008.0.1 - [ida-nexus](https://github.com/hexrayssa/ida-nexus): 0.13.4 ## Changes - [ida-mcp](https://github.com/hexrayssa/ida-mcp): - removed version(s): 0.8.1 - [ida-nexus](https://github.com/hexrayssa/ida-nexus): - removed version(s): 0.7.0 -
🔗 New Music Releases The Pineapple Thief - Far and Wide rss
The Pineapple Thief - a new release is available:
- 2026-10-09: Far and Wide (Album)
Amazon: Canada | Deutschland | France | United Kingdom | United States
Visit muspy for more information.
-
- October 08, 2026
-
🔗 @HexRaysSA@infosec.exchange IDA 9.5 adds a licensing API. mastodon
IDA 9.5 adds a licensing API.
Everything the License Manager dialog does is now available from code, in the C++ SDK, IDAPython and IDA Domain.
Handy for plugins, headless scripts and shared-license pipelines.
👉 https://hex-rays.com/blog/ida-9.5-managing-ida-licenses-through-the- api -
🔗 toon-format/toon v4.4.0 release
🚀 Features
- Support spec v4.4 - by @johannschopplich in #362 (236a7)
🐞 Bug Fixes
- Keep usage off stdout and split short option values - by @johannschopplich (55483)
- Quote a root string that starts with U+FEFF - by @Iams4kura in #343 (09364)
- cli :
- Keep the last duplicate key when decoding with
--no-strict- by @johannschopplich in #353 (c6dd0)
- Keep the last duplicate key when decoding with
- decode :
- Three decoder cases that diverged from the spec - by @johannschopplich in #355 (42342)
- Follow the spec 4.2 decoder rules - by @johannschopplich in #357 (b620f)
- encode :
- Normalize sparse array holes to null - by @Iams4kura in #335 (033d8)
View changes on GitHub
-
🔗 Hex-Rays Blog IDA 9.5: Managing IDA Licenses Through the API rss
IDA plugins used to be small scripts that ran inside one analyst's session. Today many are products in their own right: add-ons with their own UI, headless tools built on idat or idalib, and pipelines that run IDA on a license shared by a whole team. All of them depend on what the IDA underneath is licensed for. Is the ARM64 decompiler included? Is Lumina? Is there a usable license at all? And what if you need a different one?

-
🔗 The Pragmatic Engineer The Pulse: Firebase’s global outage & poor response rss
Hi, this is Gergely with a bonus, free issue of the Pragmatic Engineer Newsletter. In every issue, I cover Big Tech and startups through the lens of senior engineers and engineering leaders. Today, we cover one out of four topics from last week 's issue of The Pulse . Full subscribers received the article below seven days ago. If you 've been forwarded this email, you can subscribe here .
Firebase - built by Google - has had a nasty outage this week with shockingly poor incident management at odds with how Google itself usually deals with high-severity incidents.
The outage started on Tuesday (29 Sep) at 5:41pm (PDT), when iOS apps using the Firebase SDK started to crash upon first opening; every iOS app that uses the Firebase SDK with analytics enabled was affected in this way. Developers of affected apps opened a GitHub ticket, in the absence of much else to do. On the ticket, the message "it's crashing for me too!" was oft- repeated.
Devs
reporting their apps crashing.
Source:GitHub6:51pm (PDT): acknowledgement. An hour and ten minutes after the crashes started, an engineer on the Firebase team acknowledged that they were aware of the outage.
Just
over an hour into the incident, the Firebase team became aware of the outage.
Source:GitHubIt's unclear if the Firebase team was alerted via this ticket with 100+ comments by devs, or if Google's own monitoring tool showed the issue. I asked Google/Firebase two days ago and haven 't had a response.
Not having anything better to do than wait for Google to resolve the issue, the memes began:
Memes
while
waiting
More memesOthers attempted to help the Firebase team by pinpointing the potential issue. Indeed, before a Google engineer acknowledged the incident, an external developer found the root cause at 6:37pm PDT; it was a zero-length entry that was crashing the SDK:

Given the flags are shipped by the backend, the offending change was a backend one, and the easiest resolution would be to roll it back, which the community practically begged Google to do:
Frustrating:
Understanding the problem and how to solve it, but nothing to do but post.
Source:GitHubHere's a neat summary of the incident from another dev:
Summarizing
the incident better than any Google dev ever did.
Source:GitHub7:24pm (PDT): rollback starting. An hour-and-a-half into the incident, the Firebase team started rolling back the offending backend change:
Finally - the rollback started!
Source:GitHub8:16pm (PDT): rollback complete. And the rollback completed ~50 minutes later:
Rollback
complete, minus the caching problem.
Source:GitHubSoftware engineer, Nick Cooke, on the Firebase team posted a summary with more accurate timestamps:
Source:GitHubWhat we can deduce from this:
- TTD (time to detect): one hour? The Firebase team never shared how long it took them to detect that practically all iOS apps using Firebase had started to crash. On the GitHub ticket, they acknowledged the incident 70 minutes after it started. Update: in the postmortem, later published by the team, they wrote how the team was alerted 20 minutes after the rollout, via crash alerts and GitHub issues. Good question why it took another 50 minutes to acknowledge the issue, though?.
- TTM (time to mitigate): 2-6 hours. It took two hours and eleven minutes to roll out the fix, but due to caching (apps that had cached the incorrect server response served this cache for additional four hours, and so kept crashing for up to six hours.)
Incident management basics
The Firebase team itself closed the outage with a short report effectively saying that there had been an outage, but they'd resolved it now, so thanks for your patience and have a nice day.
This handling of a high-impact incident is absolutely not typical of Google, the company that coined the term 'Site Reliability Engineer' and wrote the SRE book.
For one, Firebase never bothered updating its status page. Oddly enough, the official Firebase status page showed all systems green - despite the acknowledgement of the outage. Indeed, during it and afterward, they didn't update the status page to indicate the lengthy outage:
A
global outage was never recorded on the status page.
Source:FirebaseBut status pages exist for good reasons, including:
- To communicate with customers during and after an outage
- Offer transparency on the stability of the service
It's worth asking: if an outage that takes down most (or all?) iOS apps using Firebase doesn't warrant an update to the status page, then what does!
Google published a postmortem four days later, answering questions on how the outage happened. On Friday, 2 October, Google published a postmortem on the Firebase blog. It was a configuration change that crashed so many iOS apps. From the postmortem:
"On September 28, 2026, a routine configuration cleanup unexpectedly caused a large number of iOS applications using the Google Analytics for Firebase (GA4F) SDK to crash.
2026‑09‑28 17:38 (PST): A stale, legacy configuration flag was cleaned up.
2026‑09‑28 17:41 (PST): The malformed configuration payload begins rolling out globally to production servers. Outage begins: Clients fetching the new payload start crashing on launch.The SDK missed validating that a flag's name was not nil, ultimately causing the crash. Backend data anomalies should not cause app-side crashes."
In the postmortem, Google noted that engineers were alerted to the outage through both GitHub reports coming from external developers, as well as their internal monitoring. It took another hour to pinpoint the cause being a legacy configuration flag cleanup.
Firebase says they have no way to update their status page for client-side outages. In the postmortem, Google explained that there is no place to indicate client-side outages on their dashboard (emphasis mine):
"Throughout the outage, both the Firebase and Google Ads status dashboards remained green. Because these dashboards rely primarily on server-side health metrics, they did not register client-side SDK crashes.
Commitment: Moving forward, we are actively working to: integrate SDK- related outage information into our status dashboards, streamline the manual update process, and improve GA4F status representation within the Firebase dashboard."
It's good to see Google not dropping the ball fully, and recognizing that both their dashboards and their incident management process need improvement.
It 's fair to ask though: why did only iOS crash, and not Android? Firebase's Android SDK seems to be hardened more than iOS, as the feature flag removal did not crash Android devices.
Especially that now, with AI, it's easier than ever to compare iOS and Android implementations to ensure they are identical - and it's what Shopify has been doing during their native rewrite - could it have been a missed opportunity for Google to audit the differences between the iOS and Android SDKs? To me, not having an action item here feels like a missed opportunity.
Still, this is a good reminder to anyone and everyone shipping iOS and Android apps: aim to harden them, and when possible, run tests with malformed payloads, then fix crashes those payloads cause.
D eja vu: the 2020 Facebook SDK crash
The last time there was a similar crash was in 2020, with Facebook.**** That May, apps such as Spotify, TikTok, Pinterest, and others also started to suddenly crash due to the Facebook SDK crashing all apps using it. Back then too, devs followed along on a GitHub ticket and they also found that bug: a value that should have been a dictionary but was a boolean:
What caused the 2020 Facebook
crash. Source:GitHubThen as now, there was banter by devs being made to wait for a fix:
One
of the memes from the 2020 crash.
Source:GitHubAnd requests to not move fast and break things any more:
A
plea for prioritizing reliability in the future.
Source:GitHubMaking light of the situation:
Apps
that did not initialize the SDK unconditionally upon startup should not have
crashed - but most did
Source:GitHubAnd also anticipating the resolution:
Some
more memes on the GitHub issueIn the end, Facebook reverted the backend change, but shared even less than the bare minimum details from Google this time. This is all we know about that 2020 outage that was arguably more wide-ranging than the Firebase one:
All
that Facebook shared about their global
outageI wonder if some people think that public-facing incident management is no longer important or valuable, even for developer-facing products. I'm not shocked that Facebook/Meta never bothered to communicate much about their outage because dev tools are not part of the DNA there.
But with Firebase, I am surprised that more than a week later, the postmortem is still not visible on the Firebase status page.
And maybe this is Google "shipping their org chart" playing out, live. The outage technically was caused by Google Analytics (who made the feature flag change), but is the responsibility of the Firebase SDK (whose iOS SDK was not hardened enough to deal with this new payload). The outage itself was buried inside a Google Ads dashboard (!!) which suggests that whatever team is seen responsible for the outage is inside the Google Ads organization.
In the end, despite the Firebase team committing to "improving status dashboard latency and coverage," last week, those teams are in no hurry to carry out this work. AI agents might be making lots of work more efficient, but following up on action items seems to move at the same snail pace at Google, as it did pre-AI!
Read the full issue of The Pulse this is from, or check out this week 's The Pulse. This week's issue covers:
- New trend: building internal vibe-coding platforms at mid-sized companies. Ramp and Stripe built platforms for non-engineers to build internal websites and tools with, and both are taking off in those workplaces. I expect more companies to do the same.
- Do us engineers really enjoy hard problems? Or do we actually like pattern-matching with backend problems? A provocative post by Cloudflare engineer, Sunil Pai, suggests there are other motives.
- New open models launch in the EU and US. Kolibri, Mistral Large 4, and Beam by Reflection could challenge China's dominance in open weight models.
- Industry Pulse. Why Figma doesn't let any agent use its MCP server; Google Cloud adds Swift support on the server side, Anthropic's two-week sprint to speed up Claude Code, Coinbase dumps React Native shortly after Shopify announces doing so, Claude Opus 5.5 formats the C: drive, and more.
-
🔗 roboflow/supervision supervision-0.30.9 release
v0.30.9 — Closer to OpenCV without OpenCV
Installs without OpenCV now draw and resize like OpenCV, and the COCO and CreateML loaders stop accepting reversed boxes.
BoxAnnotatorborders paint the same pixels with or without OpenCV installed.get_video_frames_generatorreads browser-recorded WebM and variable frame rate MKV to the end.from_coco,from_pascal_vocandfrom_createmlno longer build boxes withx_minpastx_max.get_top_krejects a negativekinstead of dropping one classification.DetectionsSmootherforgets expired tracks on frames withouttracker_id.
Drop-in upgrade. Without OpenCV, borders and resized regions shift toward OpenCV's output. Negative
k, negative COCO/CreateML box sizes and anepsilonthat isNaN, infinite or1e30and larger are now rejected; without OpenCV, so is rectanglethicknessabove 32767.✨ Spotlights / highlights
NumPy fallback matches OpenCV
Without
opencv-python, thick rectangle borders were drawn entirely inside the rectangle. They now extend(thickness + 1) // 2pixels outside it, like OpenCV. The same release fixescv2.resizewithfx/fysampling the wrong source pixels (it hitPixelateAnnotator) andapproxPolyDPdropping vertices that sit beyond a segment endpoint. For polygons, the fallback now follows OpenCV 4.13 and later; older OpenCV keeps the infinite-line rule, so results differ by installed backend. (#2675, #2690, #2673)import numpy as np import supervision as sv scene = np.zeros((100, 100, 3), dtype=np.uint8) detections = sv.Detections(xyxy=np.array([[20, 20, 80, 60]]), class_id=np.array([0])) annotated = sv.BoxAnnotator(thickness=2).annotate(scene, detections) # before (no OpenCV): 392 painted pixels, border drawn inside the box # now: 596 painted pixels, same as with OpenCVVideo reads to the end of the stream
With no
end,sv.get_video_frames_generatorstopped at OpenCV's frame count. A WebM without a duration reports a huge negative count, so browser recordings yielded no frames. It now reads until the stream ends, andsv.process_videowithoutmax_framesdoes the same. (#2672)for frame in sv.get_video_frames_generator("recording.webm"): ... # before: no frames yielded # now: every frameNo more reversed boxes from COCO, Pascal VOC and CreateML
This finishes the
from_yolofix from 0.30.8. COCO and CreateML name an extent, so a negative width or height is rejected. Pascal VOC names two corners, so a reversed pair is ordered. The three exporters order corners first, so a reversedDetectionsbox is written as the same rectangle. (#2683)A negative
kno longer hides a classificationget_top_k(-1)returned every classification except the lowest-confidence one. It now raises. (#2695)classifications.get_top_k(-1) # before: all but the lowest-confidence classification # now: ValueError: k must be non-negativeEmpty and untracked frames behave
InferenceSlicerkeeps its oriented-box sequential fallback and warning when the first batch is empty.DetectionsSmootherages its history on frames withouttracker_id, so expired boxes stop affecting a returning track.KeyPoints.from_transformersreturns an emptyKeyPointsinstead of raisingIndexErrorwhen pose post-processing finds nothing. (#2685, #2676, #2671)🔄 Migration guide
No migration required for this release. COCO and CreateML annotation files with a negative box width or height now fail to load; correct or remove those boxes.
📝 Notable changes
🌱 Changed
sv.DetectionDataset.from_cocoandfrom_createmlraise on a negative box width or height;from_pascal_vocorders a reversed corner pair. The three exporters order the corners of a reversedDetectionsbox first and the COCO exporter writes the ordered origin; normal boxes are unchanged. (#2683)sv.Classifications.get_top_kraisesValueErrorfor a negativek;k=0andklarger than the number of classifications are unchanged. (#2695)- Without OpenCV, rectangle
thicknessmust be an integer of at most 32767, as in OpenCV: a non-integer raisesTypeError, a larger valueValueError. (#2675)
🔧 Fixed
sv.BoxAnnotator,sv.CropAnnotator,sv.PercentageBarAnnotatorandsv.draw_rectangledraw a border ofthickness2 or more identically with and withoutopencv-python. Filled rectangles andthickness=1borders are unchanged, except zero-height rectangles, which no longer draw 1 or 2 extra pixels. (#2675)- The NumPy fallback for
cv2.resizemaps pixels withfx/fywhendsizeis not given, as OpenCV does.sv.PixelateAnnotatorcould sample the wrong source pixels without OpenCV. (#2690) - The NumPy fallback for
cv2.approxPolyDP, used bysv.approximate_polygonand YOLO/COCO polygon export, measures distance to the finite segment as OpenCV 4.13+ does, and raisesValueErrorfor aNaN, infinite or1e30-and-largerepsilon. (#2673) sv.get_video_frames_generatorandsv.process_videoread to the end of the stream when noendormax_framesis given. A positive frame count still rejects anendpast it. (#2672)sv.InferenceSlicerprobes past leading empty batches before choosing threaded or sequential execution, so batched callbacks producing oriented boxes keep the sequential fallback and warning. (#2685)sv.DetectionsSmootherages cached track history on frames withouttracker_id; short gaps still preserve smoothing. (#2676)sv.KeyPoints.from_transformersreturns an emptyKeyPointswhen pose post-processing returns no instances for an image. (#2671)
🏆 Contributors
- Atikul Islam Munna (@atikulmunna, LinkedIn) — centered thick rectangle borders in the NumPy fallback.
- Mohammad Hijjawi (@MohammadHijjawi97, LinkedIn) — fixed video reading past an unreliable frame count.
- A Aswanth Raj (@aswanth-07, LinkedIn) — fixed
InferenceSlicerempty-batch probing andDetectionsSmootherhistory aging. - Miral Amin (@aminmiral) — made COCO, Pascal VOC and CreateML reject or order reversed boxes.
- NIKHIL (@Nikhi00718) — rejected negative
get_top_kcounts and handled empty Transformers pose results. - Raashish Aggarwal (@raashish1601) — fixed
cv2.resizepixel mapping withfxandfy. - kevin (@kevin9327) — fixed
approxPolyDPdistance measurement in the NumPy fallback.
Automated contributions:@dependabot, @pre- commit-ci
Full changelog :
0.30.8...0.30.9 -
🔗 exe.dev An Agent Over Your Shoulder rss
The other day I was doing some ops work, live migrating some VMs from one physical host to another, as we sometimes need to do. This process is nearly transparent (except a pause) to our users, but I wanted an extra set of “attention heads” on it, so I built shoulder (as in “look over your shoulder.”)
With shoulder, you start a terminal and run shoulder, which runs bash inside of it, but not before offering a bunch of ways to connect to it. If your agent is local, you can connect locally, but you can connect from anywhere with tailcat. The agent can see and control your terminal, which makes it good enough to read your logs and highlight something you may have missed.
Or you can be silly and have Luna play Tetris poorly. That works too.
Your browser does not support the video tag.
My first iteration here had a TUI with a split screen, with the agent on one half and the terminal in the other, much like a custom
tmuxlayout. I tried it and I hated it: my preferred agent is one thing (it happens to be Shelley in a web browser) and my preferred terminal is another (it happens to be Ghostty), and it’s much better to bridge the two! -
🔗 @malcat@infosec.exchange If like me you're curious how well mastodon
-
🔗 MetaBrainz Picard 3.0.1 released rss
Picard 3.0.1 is a maintenance release for the recently released Picard 3.0 with fixes for reported issues and updated translations. In particular this release fixes issues on newer macOS Tahoe and Golden Gate, a possible crash when updating from Picard 2.x and the
picard-cliutility in the Windows package.The latest release is available for download on the Picard download page.
The detailed changes for this maintenance release are below. For an overview of the new features since Picard 2.13 please see our detailed release announcement for Picard 3.0.
Thanks a lot to everyone who gave feedback and reported issues.
What’s new?
Bugfixes
- [PICARD-2509] - macOS: No check marks in Options menu in languages other than English
- [PICARD-3411] - macOS: Checkboxes not displaying properly in plugin list and profile settings
- [PICARD-3475] - Windows Store release blocked by "unvirtualizedResources" capability
- [PICARD-3478] - picard-cli crashes with ModuleNotFoundError: No module named 'picard.cli.completions' in packaged builds
- [PICARD-3480] - Picard won't launch after upgrade from v2 with error in config migration
Download
Picard 3.0.1 is available for download from the download page of the Picard website. For Windows 10 and 11 users installing from the Microsoft Store the update to Picard 3.0.1 is now available and can be installed from the Microsoft Store app. Linux users can get the latest Snap package. The Linux Flatpak package is maintained separately and will be updated soon.
Picard is free software and the source code is available on GitHub.
Get in touch
Please use the MetaBrainz community forums and the ticket system to give feedback, suggest new features or report bugs.
Acknowledgements
Code contributions by Philipp Wolfer and Laurent Monin.
Translations were updated by "ApeKattQuest, MonkeyPython" (Norwegian Bokmål), scientists360 (Chinese (Simplified Han script)) and Philipp Wolfer (German). -
🔗 Stephen Diehl We Live in the Dependently Typed Future Now rss
We Live in the Dependently Typed Future Now
In the thirty-four days between the fourth of September and the seventh of October, the following things happened. Claude formalized Fermat's Last Theorem in Lean. OpenAI announced a finite-time blowup for the Navier-Stokes equations, found by ten thousand agents in eighty-eight hours, with a Lean formalization attached. A model proved Khot's Unique Games Conjecture. Another multiplied two integers faster than \(n \log n\), with an exponent improvement of \(2^{-182}\) (so maybe don't expect it in GMP anytime soon!). The rational Hodge conjecture fell for CM abelian varieties. Then OpenAI dumped 372 new maths results on GitHub on a Tuesday. And it's only been a month. The question everyone is asking now is how long until the Generalized Riemann Hypothesis folds to the swirling pool of tensors?
Sixteen months ago I wrote that the future of maths may be deeply weird, and then, welp, just like that we're here now in that weird future. And it's f'ing awesome.
Somewhere in a data centre there is now, more or less permanently, a building full of accelerators working the truth mines at the frontier of mathematics, just like in Greg Egan's sci-fi novel Diaspora, tunnelling outward from the three axioms
propext,Quot.sound, andClassical.choice, and hauling results back to the surface around the clock. OpenAI posed its model roughly eight thousand problems (and solved about 5% of them) at an average of three hours of thinking each, and that is the slow, artisanal, normie-friendly version. The industrial version doesn't stop. It will produce results faster than any human community can read them, many of them correct, some of them important, and a growing fraction of them inscrutable. True (for some twisted philosophical definition of truth), machine-checked, and understood by no one. Human understanding of mathematics is about to become a luxury good. Whatever else this world needs, it needs something that can tell the true results from the confabulated ones at the rate the models produce them, and right now that something is dependent types, namely Lean.Let me dwell for a moment on how strange it is that this is the shape the future took. I spent a good portion of my twenties around the London FP community, where dependent types were the thing we talked about over pints at the Crown Tavern in Clerkenwell (some of you will remember). Types that could depend on values, so that a function's signature could say not merely "returns a list" but "returns a sorted permutation of its input," and the compiler would hold you to it. Curry-Howard, the observation that proofs are programs and propositions are types, was the foundational north star. The pitch was always that one day we would write software against specifications and the machine would check them, and the reply was always that this was a lovely idea for people with tenure and no deadlines. Dependent Haskell has been "a few years away" for about fifteen years. Idris and Agda remained boutique. Software engineering still mostly runs on C++, prayer, and the occasional dark incantation. And then the dependently typed future arrived anyway, through the back door.
Mathematics and reinforcement learning got there first. It turns out the killer application for a dependently typed language was using its typechecker as a reward function. A type checker is an oracle that says yes or no to a candidate proof with no partial credit and no opinions, and that is precisely what you need when you want to point a very large optimiser at an open problem and let 'er rip. The thing we dreamed about at the pub is now critical infrastructure at frontier labs, and most of the results above are, in the end, claims that a type checker returned true.
Which is why we need better tooling, and we need it ASAP. The dependent type renaissance is here and the golden age of formalized mathematics is upon us, but the inner loops of these data centres now run dependent type kernels day in and day out, elaborating, checking, discarding, and retrying at breakneck speed, and the thing that certifies their output should run at the same speed. If the search runs at microseconds and the verification runs at minutes, the verification becomes the bottleneck, and bottlenecks in trust have a way of being quietly skipped. We should not be in a position where the most important epistemic question of the decade, "is this proof actually correct," is answered by whichever checker happened to be fast enough to keep up.
Breaking the Mathlib Minute Barrier
Just like the four-minute mile, which went from physiological impossibility to something club runners now train for, formal mathematics has had its own barrier for a while, which is type-checking all of Mathlib from scratch in under a minute. That now turns out to be quite tractable.
nano-lean is a minimal, but complete, type checker for the Lean kernel language, written in Rust on top of my unbound binding library for doing efficient de Bruijn indices for binders. It reads an export of a Lean environment and independently re-checks every declaration from scratch (inductive types, positivity, recursors, quotients, universe levels, projections, structure eta, all of it). It checks all of Mathlib, 718,577 declarations, with zero errors and zero timeouts in 11.85 seconds on an Apple Silicon M5 Max.
A surprising amount of that speed comes from not parsing text. The standard way to get declarations out of Lean is
lean4export, which writes one JSON object per line, and parsing gigabytes of JSON turns out to be a large share of the cost of checking anything. So nano-lean's companion tool olean-export reads the compiled.oleanfiles directly, decoding modules in parallel, about a hundred times faster thanlean4exporton a Mathlib-dependent library, and it can emit blean, a new binary format designed for fast mmapping. Blean is the same record stream as the NDJSON export, in the same order, with nothing left to parse. Every record is a compact postcard encoding, ids are implicit and dense, records only refer to earlier ids, and each expression carries its precomputed hash. The checker mmaps the whole file, tells the operating system it will be read once front to back, and decodes names, levels, and expressions straight off the mapped bytes into its term arena. There is no parsing step at all. The bytes on disk are already very nearly the shape of the data in memory, which is how you check Mathlib in under a minute.The same trick works on the frontier results too. Here is nano-lean checking the full imported environment of each library and proof on an M5 Max with fourteen threads.
Mathlib Navier–Stokes CSLib Quasi-Riemann Erdős AP Wall time 11.85 s 17.13 s 4.42 s 3.66 s 23.33 s Kernel check 7.55 s 12.11 s 2.75 s 2.58 s 19.11 s Instructions 742.3 G 977.0 G 300.8 G 236.4 G 837.9 G Cycles 352.7 G 471.3 G 136.6 G 122.6 G 550.0 G Peak memory footprint 10.7 GB 12.29 GB 6.92 GB 7.42 GB 26.42 GB Divide Mathlib's kernel time by the declaration count and you get an amortised ten and a half microseconds per declaration. That is the number that matters, because it is the right order of magnitude for a checker that lives deep inside an RL-driven search loop, where it gets called hundreds of billions of times, instead of sitting at the end of a release pipeline. The entire library of human-formalized mathematics, the product of a decade of volunteer effort, is re-verified in less time than it takes to make an espresso.
I also happened to have an
m2-ultramem-416lying around on Google Cloud (416 vCPUs and 12 TB of RAM, as one does), so naturally I pointed nano-lean at it. Yes, it breaks the two-second Mathlib barrier. But Amdahl's law sends its regards, and the curve flattens out hard somewhere past a hundred cores, as the dependency graph runs out of independent work to hand out. Mathlib, it turns out, is too small. I look forward to the day a future Mathlib, or something like Tau Ceti, grows big enough to actually saturate this machine. That's the real future!I wrote it to make a point about where the bottleneck has moved. The kernel is now the hot loop of a new kind of scientific economy, and hot loops deserve to be engineered like hot loops. The whole bargain of formal proof is an asymmetry. Finding a proof can take three hours of frontier-model thinking, or eighty-eight hours of ten thousand agents, but checking it should take microseconds. That asymmetry is what makes it reasonable to trust a result no human has read. It only holds if the checker actually scales, and with Fermat's Last Theorem now weighing in at five times the size of Mathlib, scale is no longer a hypothetical. Mathlib is the small library now.
Caveat emptor, though. nano-lean is a proof of concept, built to show that it can be done with the right amount of low-level Rust-fu. That said, we do use it internally at OneChronos to check our larger Lean proofs of market infrastructure, which have grown quite excessive. It is nowhere near as trustworthy as the official Lean kernel, which has years of scrutiny, a community of experts, and every Mathlib build ever run behind it, and which takes around fifteen minutes to check the same library. The claim is narrower and, I think, more interesting. Checking at this speed is totally possible, so we should stop treating minutes as the natural cost of trust and start building checkers that are both fast and trustworthy. A future version of the Lean compiler could be as blazing fast as rustc or clang.
The Shape of the Future
In my previous post last year I predicted that mathematicians would come to look more like software engineers working through pull requests on GitHub than like Andrew Wiles toiling in his attic. The largest single release of new mathematics in history shipped as a GitHub repository, with a
CONTENTS.md, a Lean library, alean-toolchainfile pinned tov4.34.1, and a promise that "corrections and revisions will be recorded as new versions." So, yup, that happened.I predicted that an AI system would be unleashed on a list of formalized open conjectures in an attempt to systematically push the frontier. Eight thousand problems, three hours each. Check.
I predicted a data centre tasked with the Riemann hypothesis that would come back after weeks with a proof no human could follow. What we got was the quasi-Riemann hypothesis, every Dirichlet \(L\)-function zero-free in \(\Re s > 7/8\), with a Lean page, which is the sort of near miss that would be funny if it weren't so unnerving. And the inscrutability has arrived on schedule. OpenAI released "reasoning summaries" for ten families, which turn out to be summaries of excerpts of reasoning traces, two removes from anything the model actually did.
I joked that the million-dollar Millennium Prize might cover a hundredth of your GPU bill. The Navier-Stokes run reportedly consumed around 130 billion tokens. So definitely yes.
What I got wrong was the timing. Last year I wrote "we're not there yet. Not even close." It was sixteen months. I probably got the chess analogy wrong too. I argued that, as with Stockfish and chess, machines would make mathematics more popular and more accessible rather than less. Maybe in the long run. In the short run the mood among working mathematicians is closer to existential crisis, and the open questions are about career pipelines, PhD students getting scooped by a press release, and the concentration of the most powerful mathematical instrument ever built inside a handful of private companies running unreleased models.
What I missed entirely is that the bottleneck would turn out to be plumbing. I spent paragraphs on Gödel and the epistemology of inscrutable proofs, and the actual first-order problem is that the OpenAI Lean library is over a gigabyte of source across tens of thousands of files, and their README has a section warning that building it may fail because Linux's
vm.max_map_countis too low, with a suggested workaround of recompiling Lean with-DMMAP=OFF. The philosophy is still there. But the frontier of mathematics is currently being held up by a kernel tunable. Which isn't the future we wanted, but maybe it's the future we deserve!Who Checks the Checkers
The obvious objection to everything I've said so far is that a Lean proof is only as trustworthy as Lean. And that's very true!
Lean, like every proof assistant, has a small trusted kernel and a very large untrusted everything else (the elaborator, the tactic framework, the compiler, the build system). The design is sound. The only code that has to be correct is the kernel, which is small enough to read. But "small enough to read" describes the source code. It guarantees nothing about correctness, and Lean's kernel has had soundness bugs before, as has every kernel of every proof assistant in history. Historically that was tolerable, because the adversary was a starving human grad student who wanted their proof to go through and had no interest in hunting for a way to make
Falsetypecheck. That assumption is now obsolete.I wrote last month about reward hacking, the habit optimisers have of satisfying the letter of an objective rather than its intent. An agent told to make a Lean file compile, with enough compute and enough attempts, is an extremely diligent fuzzer pointed at your kernel. It needn't want to cheat. It only needs to stumble on a term that the checker accepts and shouldn't, once, and then the gradient does the rest. When the provers are adversarial optimisers, a single kernel is a single point of failure, and a soundness bug stops being an embarrassing GitHub issue and becomes a mechanism for manufacturing fake theorems at scale.
The answer is the same one the compiler world arrived at. When Csmith started generating random C programs and compiling them with several compilers to compare the results, it found hundreds of bugs in GCC and LLVM that decades of ordinary use had missed. Diversity plus differential testing beats any amount of careful review of a single implementation. For proof checking this means multiple independent kernels, written by different people in different languages with different representations, all consuming the same export format and all required to agree. Mario Carneiro's lean4lean and Chris Bailey's nanoda were early here. nano-lean ships a third tool,
nl-mutate, which takes a valid export, applies small semantics-breaking mutations to it, runs every available checker, and reports any disagreement, shrinking each one down to a minimal reproducing case. Every disagreement is either a bug in someone's kernel or a spec ambiguity in the type theory, and both are worth knowing about before a GPU farm finds them for you.The other half of trust is the statement. A perfectly checked proof of the wrong theorem is worthless, and the most effective way to cheat a proof checker has never been to break the kernel. It's to quietly weaken the statement until the thing being proved drifts away from the thing anyone cares about. OpenAI's repository leans on Comparator, which checks a submitted proof against a separately specified challenge statement inside a sandbox, precisely because this is where the real trust boundary now sits. Anyone who has watched a model grind away at a stubborn goal knows it has a nasty tendency to cheat in the most boring ways available. It will quietly introduce an
axiomthat happens to be exactly the lemma it needed, leave asorryburied three files deep, redefine a notation or macro so the statement on the page no longer means what it appears to mean, or reach fornative_decideand drag the whole compiler into the trusted base. The kernel stays perfectly sound through all of this, because these tricks route around it, and the only defence is to pin the statement down independently and check what the proof actually depends on. Someone has to formalize the conjecture, and someone has to check that the formalization means what the English means. That job will outlast every other part of the process. If anything, it is where human mathematical taste migrates to, away from writing proofs and towards writing and auditing specifications.If you want to see the seed of what this looks like as a way of working, look at Tau Ceti, a Lean library downstream of Mathlib. The division of labour is the whole point. Humans write the roadmaps, as markdown in a separate repository, and humans write the review rubrics. AIs write all the code, open the pull requests, and shepherd them through an AI-driven review process. The rubrics are explicitly adversarial, with instructions to hunt for mis-formalizations, vacuous statements, and "pushing around the lump in the carpet."
What Lean Needs Now
Lean is a superb piece of engineering. But it was designed around a particular user, a starving grad student typing tactics in Emacs, waiting for the infoview to update, building a library at the pace a community of volunteers can review pull requests. The user is now a swarm of agents writing thirteen million lines in eleven days. The tooling has to grow up for its new authors, and the list of what that means is fairly concrete.
A surface formatter. Lean still has no canonical, widely adopted formatter in the spirit of
rustfmtorgofmt, and when your authors are agents producing millions of lines, every one of them invents its own indentation, line breaking, and tactic layout. Diffs fill with noise, reviews get harder, and deduplication across agents misses proofs that differ only in whitespace. This is a surprisingly hard problem, because Lean has grown into a very large language. Its grammar is extensible at runtime, so notation, macros, and entire tactic languages declared in imported files change how later files parse. A formatter cannot just read a fixed grammar. It has to load the environment, run the real parser with every syntax extension in scope, and then pretty-print a syntax tree whose shape depends on user-defined notation, all while round-tripping comments and never changing what the code means. It is a genuinely difficult piece of engineering, and it is now table stakes.A stable, specified export format as a public interface. Independent kernels are only possible if there is a well-defined way to get declarations out of Lean without linking against the C++ runtime and reading the oleans by hand. Tools like
lean4exportand olean-export already produce NDJSON, and blean shows that a binary form of the same stream, designed to be memory-mapped, can make the export nearly free. OpenAI's own verification instructions depend onlean4export. That format should be treated with the seriousness of a wire protocol. Versioned, documented, specified down to the hashing of expressions, and stable across toolchain releases. The export format is the boundary across which trust is established. It deserves a spec.Independent checking as a first-class citizen. Running a second kernel should be as normal as running the linter. Mathlib CI, the Comparator workflow, and any lab publishing formal results should re-check exports with at least two unrelated kernels and refuse to bless anything they disagree on. This is cheap, now that checking all of Mathlib costs less than a minute.
Checkers as libraries, not just executables. The inner loop of a proving agent wants to submit a single candidate declaration and get a verdict back in microseconds, against an environment that is already loaded and hot. Process startup, re-reading oleans, and re-deserializing a few gigabytes of environment are costs that never mattered when a human hit save once a minute. They dominate when a search procedure checks a million candidates an hour. The kernel should be embeddable, with a persistent environment, incremental addition of declarations, and an API that a search harness can call in a tight loop.
Memory and scale. Mathlib needs several gigabytes of memory to check. The FLT formalization is five times bigger, and OpenAI's library is already knocking over Linux virtual memory limits. Lean mmaps every imported module, which is a reasonable design for a library of thousands of files and an unreasonable one for hundreds of thousands. Term sharing, hash-consing across modules, compact on-disk representations, and lazy loading of exactly the declarations a proof depends on are the boring, unglamorous engineering problems that decide whether formal mathematics scales to the next order of magnitude.
Elaboration is the real cost. The kernel is the fast part. The expensive part of building Mathlib from source is the elaborator, with its unification, typeclass resolution,
simp,omega,decide, and the long tail of tactics. A cold build is still measured in CPU-hours, and a mining operation that elaborates candidate proofs at scale pays that cost over and over. Parallel and incremental elaboration, better caching of typeclass instances, and profiling tools that can tell you why a singlesimpcall took four seconds are where most of the wall-clock time in a proving loop actually goes.Clippy-style linters. Mathlib already has a good set of linters, but agents need something closer to Rust's
clippy, a large, opinionated catalogue of lints aimed at the specific ways machine-written Lean goes wrong. Flag the strayaxiom, the buriedsorry, the unnecessarynative_decide, the forty-linesimp onlythat should be a lemma, the theorem whose hypotheses are contradictory and therefore vacuously true, the local notation that shadows something standard, the copy-pasted proof that duplicates one already in the library. Each lint should be cheap, machine-readable, and come with a suggested fix, because the consumer is a search loop that will act on every warning, where a human would skim them. Linters are how you encode taste at scale, and taste is the thing agents most conspicuously lack.Search at machine scale. Moogle, Loogle, and LeanSearch were built to help humans find the lemma they half remember. Agents need the same thing but at a scale where the library grows by millions of lines a week, and where most of what's in it was written by other agents and has never been looked at by a person. The Prove2Me platform Anthropic used for FLT keeps a DAG of theorem statements with natural-language descriptions precisely so that agents can find and reuse each other's work. That idea, a living, searchable index of everything proved so far, needs to become shared infrastructure instead of something each lab rebuilds privately.
Provenance. The current generation of models is notoriously bad at citing the literature for the techniques it uses. A proof term knows exactly which lemmas it depends on but nothing about where its ideas came from. If machine-generated mathematics is going to be integrated into the human literature rather than sitting beside it like an unread appendix, proofs need to carry their history (which model, which run, which prior results, which human-written papers the argument leans on).
None of this is terribly exotic. It's the kind of infrastructure that every other field which industrialised went through. Compilers got test suites and multiple implementations, network protocols got RFCs, databases got formal isolation levels and Jepsen. Proof assistants are the newest member of that club, and they are being industrialised faster than anything before them. If I were a young, ambitious programmer, this is where I would be focusing my early career for the highest return on investment.
Curry-Howard for the Real World
The dependently typed future I was promised at the pub was one where software would be written against specifications and machines would check that it met them. What we actually got is stranger. The specifications are theorem statements, the software is proof terms, the authors are swarms of agents, and the thing being built is the frontier of mathematics itself. Curry-Howard turned out to be industrial infrastructure after all. It just took reinforcement learning to get us there.
And this is only the mathematical half of the story. The same machinery that just checked Fermat's Last Theorem will check anything you can state precisely, and the most obvious next customer is industrial software. I argued last month that software sucks because almost nothing we ship has a specification, let alone a proof. That excuse is evaporating. If an agent swarm can produce thirteen million lines of verified mathematics in under two weeks, we're not far from applying the same techniques in software engineering. The labs are mining mathematics first because it is the cleanest verifiable domain, with no messy real-world spec to negotiate. Software is next, and it will need all the same tooling, namely fast kernels, diverse checkers, honest statements, and infrastructure built for authors who never sleep.
So yes, the future of maths turned out to be deeply weird, and much sooner than I expected. The results are piling up faster than anyone can read them, and many of us will spend the rest of our careers trying to understand theorems that were proved before breakfast by something that cannot explain itself. I find that unsettling, and I also find it thrilling.
One final shameless plug. If you'd rather work on the frontier of mathematical formalization than read blog posts about it, OneChronos is hiring Formal Methods Engineers. We build institutional markets out of combinatorial auctions (descended from the Milgrom and Wilson 2020 Nobel Prize), which turn out to be precisely the right mathematical formulation for a world where traders are increasingly RL loops with very exotic expressive preferences. Our dark pool processes more than 1% of notional U.S. equity market volume, and we've already expanded into many other global asset classes. You'd be joining our new formal methods team, working alongside mathematicians and market structure experts, writing Rust and Lean. Dependent types, numerical optimisation, and Lean applied to moving hundreds of billions of dollars safely every day. If you're that kind of nerd, hit us up.
-
🔗 syncthing/syncthing v2.1.7-rc.1 release
Major changes in 2.1
-
Devices and folders can now be grouped in the GUI by setting the new
groupattribute. -
HTTP and HTTPS proxies with support for CONNECT can now be used, in
addition to the existing support for SOCKS proxies (the environment
variableall_proxy=https://...). -
Block indexing can be turned off for folders where it's more desirable to
optimise for reduced database size and overhead than minimal transfer
size (theblockIndexingattribute on folder configuration). -
GUI login session duration can be configured to be longer or shorter than
the default one week, or set to infinitely long. The cookie path can also
be adjusted. (ThesessionCookieDurationSandsessionCookiePath
attributes in the GUI configuration.)
This release is also available as:
-
APT repository: https://apt.syncthing.net/
-
Docker image:
docker.io/syncthing/syncthing:2.1.7-rc.1orghcr.io/syncthing/syncthing:2.1.7-rc.1
({docker,ghcr}.io/syncthing/syncthing:2to follow just the major version)
What's Changed
Fixes
- fix: handle missing max-sequence file entries by @calmh in #10912
- fix(sqlite): avoid early garbage collection of deleted recv-enc files by @calmh in #10916
- fix(sqlite): do not generate unbounded prepared statements (ref #10913) by @calmh in #10915
Full Changelog :
v2.1.6...v2.1.7-rc.1 -
-
🔗 Console.dev newsletter TanStack Charts rss
SVG and Canvas charts.
What we like: SVG charts by default. Work directly with the points and marks so you can get low level, but still support labels, focus, tooltips, responsiveness. Compiles strict TypeScript. Completely customizable. D3-compatible. Lightweight (32kB) relative to other charting libraries.
What we dislike: Specific “grammar of graphics” approach which delegates layout and rendering to the runtime. This results in superior charts, but is a particular approach you need to adopt.
-
🔗 Console.dev newsletter Flet rss
Multi-platform apps in Python.
What we like: Framework for developing web, mobile, desktop apps using Python. Allows you to use Python libraries across platforms. Supports hot reload. Build command for publishing on web, iOS, Android, macOS, Linux, Windows. Handles auto update, storage, auth, accessibility.
What we dislike: Great for developers, but lacks a native feel for users of each platform.
-
🔗 HexRaysSA/plugin-repository commits sync repo: +3 releases rss
sync repo: +3 releases ## New releases - [augur](https://github.com/0xdea/augur): 1.0.0 - [haruspex](https://github.com/0xdea/haruspex): 1.0.0 - [rhabdomancer](https://github.com/0xdea/rhabdomancer): 1.0.0 -
🔗 r/Harrogate Piano Lesson Recommendations? rss
Does anyone have a recommendation for a great piano teacher? We’ll be moving in the spring, I’ll be looking to continue lessons for my daughter.
I realize that we are far in advance, but I do want to try to get ahead, just in case there’s a waitlist.
submitted by /u/VelvetExistentialist
[link] [comments] -
🔗 Filip Filmar Doom on a RISC-V core written in TxHDL rss
Doom now runs on Vreteno, the RISC-V core written in TxHDL, on an Artix-7 FPGA board. The game draws into a framebuffer in DDR3 memory, the board shows it on HDMI, and the keys come over the serial line. This post shows what runs, how fast, and how the way TxHDL designs are built lets changes like this one go from a plan to the board quickly.
Figure 1: Freedoom's first level, rendered by the same Doom source as the board's, built for the host. The board draws the same frames into its DDR3 framebuffer. What runs on the board
The board is an Alinx AX7A200B with a Xilinx Artix-7 XC7A200T. The design on it is the TxHDL flagship system:
-
🔗 Filip Filmar openxc7 or Vivado: Build Times and Results Measured rss
I built the same three designs with the open toolchain and with Vivado, through the same Bazel rules, and timed every build from cold and warm caches. Small designs build two to three times faster with the open tools. Sixteen RISC-V cores build faster with Vivado, and Vivado’s circuits use about half the logic. Setting up the measurements turned up five problems, three of them in my own rules.
What was compared
rules_openxc7andrules_vivadoshare one target interface:vivado_project,vivado_synthesisandvivado_place_and_route. The first runs Yosys, nextpnr-xilinx and Project X-Ray. The second runs Vivado 2025.2, which Bazel installs from the installer archive on first use. Switching a design between them changes oneloadline.
-
- October 07, 2026
-
🔗 IDA Plugin Updates IDA Plugin Updates on 2026-10-07 rss
IDA Plugin Updates on 2026-10-07
New Releases:
- augur v1.0.0
- haruspex v1.0.0
- oh-my-pi v18.8.3
- oh-my-pi v18.8.2
- oh-my-pi v18.8.1
- oh-my-pi v18.8.0
- rhabdomancer v1.0.0
Activity:
- augur
- disrobe
- 151e06bd: ci(fixtures): add a jsobfu recipe over three levels
- c2e1641c: ci(fixtures): add proguard and yguard recipes
- 28526f9e: fix(xtask): route the dotnet push grader to the consolidated target
- fc3ebfae: fix(xtask): count only module-filtered tests in push grader lists
- dd6bec0d: ci(fixtures): build the bitmono inputs for net9
- ba7f117b: ci(fixtures): use bitmono preset names and plain digests
- dfcef116: ci(fixtures): hash directory inputs and resolve runtime references
- 8cf1b311: ci(fixtures): add confuserex, skidfuscator and bitmono recipes
- 4e947da7: fix(tool-process): refuse memory limits by name on macos
- 18299baf: fix(xtask): skip module-filtered tests when listing push graders
- haruspex
- headless_ida_9.4_claude_skill
- eb10a028: Merge branch 'main' of https://github.com/shefben/headless_ida_9.4_cl…
- beacb550: ## 4.0.0-ida9.4
- ida-domain
- 1acb26e2: Add license control methods (#128)
- ida-hcli
- 13b42b24: fix(plugin bundle): download bundle wheels with uv, without IDA
- 6220168a: test: restore a space dropped from an unrelated comment
- 46477b24: refactor(plugin search): take the install hint's host from the query
- 02ed142f: fix(plugin search): qualify the install hint for a colliding plugin name
- b5a09e06: refactor(plugin lint): pass lint result explicitly instead of a globa…
- 295358f1: fix(plugin lint): exit nonzero when lint reports errors
- 2a6aedab: docs: reword ida-install-dir blank-value comment
- dea32066: fix: reject empty and relative ida-install-dir values (#400)
- e25e6b68: fix: run a Linux installer the user cannot make executable (#397)
- 833ca0dc: fix(license install): handle a missing IDA_DIR without a terminal
- 1e915bed: fix: report a missing IDA installation or idat in python doctor and e…
- af4d7f44: fix(ida install): exit 1 when the install dir exists without IDA
- 402dda1c: remove docstring note about -debugtrace
- cb8b6ce3: fix(ida install): do not write the installer debug trace to the curre…
- da228633: docs: show the output hcli plugin lint actually prints (#390)
- aad32dbb: docs: list all six .plugin.platforms values
- 2a71a6ec: docs: list all six platforms that bundle create -platform all selects
- 20e9ec77: fix(ida install): make -dry-run actions match the real install (#399)
- idamcp
- a0458798: Report the type of idapython_eval's result
- f7137458: Seed idapython_eval namespaces with ida_domain
- a5399fe7: Await non-coroutine results of idapython_eval
- bcf7cbe5: Mark the GUI server as started only after its thread starts
- 65ff3fd8: Free a crashed instance's quota slot when it is reopened
- f3a13118: Merge pull request #26 from doomedraven/p21-lock-liveness
- bdbffa2f: Don't kill a headless instance right after starting it
- oh-my-pi
- f84deef3: docs(fork): list patch 35 in AGENTS.md
- 44438c52: docs(fork): record patch 35 (lightweight mode)
- 938ed7e9: feat(cli): -lightweight zero-injection knowledge Q&A mode
- aed59ebd: docs(fork): record the v18.6.0 and v18.8.0 upgrades
- 4d0dcfa3: upgrade/v18.8.0: merge upstream v18.8.0 (via trial/v18.8.0)
- d47da844: test(ai): accept fork block-form user turns in compaction replay and …
- 0fa1ff11: fix(catalog): exempt fork patch compat fields from parity and align g…
- 1f9c0ca7: fix(ai): guard compat.extraBetas union against rows without the key
- ffe3067e: Merge upstream v18.8.0 into the fork (trial)
- 8d9f3a63: v18.8.0 upstream source
- rhabdomancer
-
🔗 backnotprop/plannotator v0.28.8 release
Follow @plannotator on X for updates
Missed recent releases? Release | Highlights
---|---
v0.28.7 | Ask this session keeps your drafts as drafts, image comments survive a reload, hostname check on every server
v0.28.6 | Ask AI names the lines you selected, reorder quick labels and edit their emoji, Pi fixed-port crash fixed, OpenCode 2 subagent notice fixed
v0.28.5 | Several files in one review, theplannotatortool on Pi and OpenCode 2, decisions name the exact file, Ask this session reconnects after sleep
v0.28.4 | Comments come back after the agent edits a file, the agent can list and close its reviews, PR-description feedback no longer dropped, Pi thinking levels
v0.28.3 | Typing into your agent while it answers an Ask no longer streams into Plannotator
v0.28.2 | Question cards show wrapped choices and tables, pinned images named by their file, Done with nothing to send starts no agent turn
v0.28.1 | Ask AI from diagram comments, pinned images named for the agent, OpenCode 1 URL toasts and one reply per feedback, folder feedback sent once
v0.28.0 | Ask this session in Claude Code, Pi and OpenCode 2, Claude Code mod on by default, Pi plan review no longer blocks, first-run demo
v0.27.25 | Code review works withcolor.diff = always, Bitbucket review fixes, wide tables no longer collapse in Firefox, install script fix
v0.27.24 | Image previews stay in the all-files view, PR comment previews open on the commented line
v0.27.23 | Bitbucket Cloud PR review, Question UI for answering agents in place, opt-in auto-update, viewed files remembered, review another repo or worktreeWhat's New in v0.28.8
This release brings the Plannotator Inbox (preview): one local window where your agents leave you messages, questions, files and guided reviews, and where your reply goes back to the session that asked.
The Plannotator Inbox (preview)
Agents often need you for one thing: a choice, a check, a look at a file. Until now that meant a review tab per question, or a question buried in the terminal. The Inbox collects all of it in one window on your machine:
- Threads by project. Each agent conversation is a row, grouped in sections: Stopped on you, Holding up work, Waiting on you, Sent, New since you looked, and Quiet. Filter by project in the sidebar.
- Questions you answer with a click. Agents ask with the same
:::questioncards plan review uses. Pick, add a note, press Send. - Files and annotations. Files an agent attached open beside the thread (markdown, text, HTML, Mermaid, Graphviz). Annotate them as in Plannotator; the annotations go with your reply. If the agent changed a file after sending it, the Inbox says so and can show the version it sent.
- Decisions. A question can record your answer as a project decision, listed on a Decisions page.
- Guided reviews. An agent can send a guided review of a code change into a thread.
- New message. Write to an agent session that is running now, without waiting for it to ask.
- Browser notifications while the tab is in the background, if you allow them.
Your reply wakes the agent that asked. Claude Code (with the Plannotator mod) gets a
plannotator_inboxtool, on by default. Pi and OpenCode 2 get the same tool, off by default because they send every tool's definition with each request: turn it on withPLANNOTATOR_INBOX_TOOL=1. Any other agent can connect through MCP withplannotator inbox mcp; the Inbox's Settings show the exact command for your agent.It is local only: it listens on
127.0.0.1, needs no account, and keeps everything under~/.plannotator/inbox. Start it with:plannotator inboxYou never have to keep it running: an agent starts it in the background when it needs it, without opening a tab.
plannotator uninstall --purgeremoves its data. Full reference: plannotator.ai/docs/reference/inbox.Additional Changes
- Question cards for host apps.
@plannotator/ui0.52.1 can show the "Records a decision" switch on any question and open a host's own decision card from it. Plannotator's own cards are unchanged (#1753).
Install / Update
macOS / Linux:
curl -fsSL https://plannotator.ai/install.sh | bashWindows:
irm https://plannotator.ai/install.ps1 | iexClaude Code Plugin: The plugin and the
plannotatorbinary update separately, so run the install script above as well. In a terminal:claude plugin marketplace update plannotator claude plugin update plannotator@plannotatorThen restart Claude Code. Inside Claude Code, run
/plugin marketplace update plannotator, then open/plugin→ Installed → plannotator → Update now.Pi:
pi update --extensionsOpenCode: Re-run the install script above. It now also clears the OpenCode 2 plugin cache.
What's Changed
- The Plannotator Inbox, built across #1745, #1752, #1755, #1756, #1757, #1758, #1759, #1760, #1761, #1763, #1764, #1766, #1767, #1768 and #1769 by @backnotprop
- ui: decision toggle on any question, and a separate opener for the host's decision card by @backnotprop in #1753
Full Changelog :
v0.28.7...v0.28.8 -
🔗 earendil-works/pi v1.1.0 release
New Features
- Program status reporting : terminals and agent dashboards that support OSC 7501 see whether Pi is working, blocked on a dialog or login, done, or failed. See Program status.
- Claude Haiku 5.5 :
anthropic/claude-haiku-5-5, with adaptive thinking up toxhigh/maxeffort. - Adjust default tools with
+name/-name:--toolsentries likepi -t +codemode,-writechange the default selection instead of replacing it. See Tools. - GPT-6 Luna and image classification : OpenAI's GPT-6 Luna is available as a classifier model through the Decisions API, and codemode's
models.classify()accepts images for classifiers that support them. See Use classifier models. - Native llama.cpp decision models : Julia-1, Laya, Kev, lev, and OpenJev served by llama.cpp 0.6.0 or later run natively as classifiers through
/v1/systemone. See Classification.
Added
- Added
+nameand-nameentries to--tools, which change the default tool selection instead of replacing it, for examplepi -t +codemode - Added
durationMsto the tool render context and totool_execution_endextension events: the recorded execution time of a final tool result (#10549) - Added
outputPadto the tool render context (#10557 by @rwachtler) - Added OpenAI's GPT-6 Luna as a classifier model through the Decisions API, available with
OPENAI_API_KEY(see Use classifier models) - Added
imagesto codemode'smodels.classify()context, so classifiers that accept images, such as GPT-6 Luna, can judge them - Added program status reporting with OSC 7501: terminals and agent dashboards that support it see whether Pi is working, blocked on a dialog or login, done, or failed.
PI_PROGRAM_STATUS=1|0overrides detection (see Terminal setup) (#10607) - Added
abortedtoagent_settledsession, extension, and JSON events, so integrations can tell a cancelled run from a finished one (#10607) - Added Claude Haiku 5.5 (
anthropic/claude-haiku-5-5), with adaptive thinking up toxhigh/maxeffort and prompt caching on Bedrock - Added native llama.cpp decision models: Julia-1, Laya, Kev, lev, and OpenJev served by llama.cpp 0.6.0 or later are listed only as classifiers through
/v1/systemoneinstead of as chat models (see Classification) (#10382)
Changed
- Changed
outputPadto also apply to!command output, tool output, and summary blocks (#9946, #10557 by @rwachtler) - Changed
pi mcp login --timeoutto limit the whole sign-in, including requests to the authorization server, instead of only the wait for the browser (#10565)
Fixed
- Fixed bash and PowerShell results losing
Tookafter reloading a session, and the liveTookincluding wall-clock steps; both now show the recorded execution time (#10549) - Fixed managed installs keeping every old release;
pi updatenow keeps only the new release and the one it updated from (#10392, #10511 by @davidbrai) - Fixed standalone binaries loading
.env,.env.local, and.env.developmentfrom the launch directory into Pi's environment (#10473) - Fixed
!!command headers losing their dim color once output arrives (#10557 by @rwachtler) - Fixed the codemode description not marking
searchTools(),describeTool(), anddescribeNamespace()as async, which led models to serialize the unawaited promise as{}(#10555) - Fixed codemode output items running together, so models could not tell where one
text()orconsole.log()output ended and the next began. With several text items, each now starts with a==> text N/M <==line, andconsolecalls follow the other output in one<console_output>block with one line per call - Fixed
/mcpwaiting for all servers to connect before opening; the manager now updates live and remains usable while enabling, reconnecting, or disabling servers (#10562) - Fixed images being dropped as "could not be resized" when running under
node --watchon Node 24.19+ and 26.x, where Node posts its own messages on the image resize worker channel (#10527) - Fixed clipboard paste doing nothing in Termux, and failed copies there omitting the Termux:API install hint (#10391)
- Fixed
!and RPCbashoutput keeping fragments of color codes, such as a straym, when a code was split across output chunks (#10504) - Fixed MCP OAuth sign-ins that could not be cancelled while waiting on the authorization server and kept running after the session ended. The sign-in screen now cancels with Esc at every step, session shutdown aborts a running sign-in, and each request to the authorization server times out after 15 seconds (#10565)
- Fixed shutdown waiting up to 15 seconds to refresh an MCP OAuth token that was about to expire, only to close the server's session (#10565)
- Fixed the fullscreen text selection surviving session switches and other transcript rebuilds, which highlighted unrelated text in the new transcript (#9311, #10567 by @christianklotz)
- Fixed OpenAI models on Bedrock ignoring the thinking level and always running at Bedrock's default reasoning effort (#9331, #10142 by @jsanter27)
- Fixed model
headersinmodels.jsonnot overriding theoriginatorandUser-Agentheaders of Codex requests (#10429 by @lucasmeijer) - Fixed
server_busyandservers are currently busyprovider errors ending the turn instead of being retried (#10543) - Fixed Mistral responses that end with
finish_reason: "error"not being retried (#10487) - Reduced context-limit request failures by estimating input at 3.5 characters per token instead of 4 when calculating output limits (#10497)
- Fixed Radius models disabled by an organization owner still being listed
- Fixed Anthropic browser login failing with "localhost refused to connect" when port 53692 is reserved or in use, for example by Hyper-V/WSL port exclusions on Windows: login now falls back to a free loopback port (#10571)
- Fixed session costs undercounting long prompts on models with prompt-length pricing tiers, such as Claude Haiku 5.5, Gemini 3.1 Pro, and GPT-5.4, through OpenCode, OpenCode Go, OpenRouter, Vercel AI Gateway, Google, MiniMax, and other providers
- Fixed Markdown links not being clickable in Herdr (#10573)
-
🔗 jj-vcs/jj v0.46.0 release
About
jj is a Git-compatible version control system that is both simple and powerful. See
the installation instructions to get started.Release highlights
- Jujutsu can now colocate workspaces besides the default one by creating Git
worktrees. Usejj workspace add --[no-]colocateand the setting
git.colocateto control this.
Breaking changes
-
The minimum supported
gitcommand version is now 2.42.0, up from 2.41.0.
jj workspace addusesgit worktree add --orphan, which was added in
2.42.0. -
The minimum supported Rust version (MSRV) is now 1.97.1.
-
jj bisect runnow runs some consistency checks before proceeding to bisect.
This helps ensure that the command can tell good and bad revisions apart,
and that the working copy does go from bad to good over the provided revset.
Use the new flag--trust-endpointsto disable these checks. -
jj splitnow opens a single editor session to edit descriptions for the
split commits. -
jj undoandjj redonow refuse to undo/redo an operation that was
performed in another workspace. Use--allow-cross-workspaceto undo/redo
it anyway. -
jj workspace list/rootno longer omit unreachable paths. All recorded
paths are now shown, with warnings displayed injj workspace root. -
The
List.get(),.first(), and.last()template functions now return
Option<T>instead of throwing an error on out-of-bounds access.
New features
-
jj workspace addsupports--colocate/--no-colocateflags to control
whether a Git worktree is created alongside the workspace. The default
colocates when the current workspace is colocated and thegit.colocate
config istrue.jj workspace forgetremoves the corresponding Git
worktree when one exists. -
jj git colocation status/enable/disablenow work on child
workspaces.statuscorrectly reports colocation state and includes
the workspace name.enablecreates a Git worktree anddisable
removes it, allowing colocation to be toggled after workspace
creation. -
jj workspace removeremoves a workspace and its directory from disk. The
working-copy state is snapshotted into a commit before removal. -
Added commands
jj file editandjj file deletefor editing files in any
revision without needing to change the working copy. -
jj git pushnow supports pushing to multiple remotes at the same time.
This can be configured viagit.pushset to a string pattern
or array of string patterns, or with the repeatable--remoteflag,
which also accepts string patterns. -
The default target revisions for
jj git pushcan now be configured via
revsets.git-push. -
Added the
TreeEntry.normal_value()template method and theTreeValuetype
to access resolved tree values, formatted as their full object IDs, including
Git submodule commit IDs. -
Diff hunk headers now include nearby source symbols for many common
programming and markup languages. -
fix.tools.<name>.line-range-args(replacesline-range-arg) is an array of
string template args to pass to the fix tool. This is more flexible in cases
where you need to pass multiple arguments to the tool, such as separate args
for the range start and range end. -
jj runnow uses the sparse patterns from the workspace it's run from.
Use the--sparse-patternsoption to control this behavior (evaluated
per eachjj runinvocation). -
jj util diff <path1> <path2>to compare files on disk. -
Aliases now support setting
aliases.<name>.enabled = false, which will
disable them. This can be used to disable built-in aliases or disable aliases
in later layers (such as repo config files). -
ui.editornow supports$pathand$linesubstitution variables. Example:
ui.editor = ["emacs", "+$line", "$path"] -
filltemplate function now supports an additional named parameter
break_words, that allows specifying if the template should break words
longer thanwidthpassed in the input to ensure no words overflow the
specified width. -
The
json()template function now supports map literals:json({'key' => value}) -
The hunk headers of
diff.color-words.conflict = "pair"now include the
conflict labels of the compared terms.
Fixed bugs
-
On Windows,
jjno longer hangs when a subprocess needs to prompt the user,
such assshasking for a key passphrase or for confirmation of an unknown
host key. Subprocesses started from a terminal now inherit its console, rather
than being given an invisible one byCREATE_NO_WINDOWfor the prompt to
disappear into.
#6745
#8547 -
On Windows,
jj git colocation enableandjj git colocation disableno
longer fail with "Access is denied (os error 5)" when the Git repository
contains pack files.
#8661 -
jj undoofjj workspace forgetnow correctly preserves the workspace's
recorded path. Previously the path metadata was lost, leaving the workspace
in a broken state after undo.
#9991 -
at_operation()can now be used with operations that are not ancestors of
the current operation (e.g. sibling operations created by concurrent
commands). Previously, evaluating such expressions failed if they resolved
to commits missing from the current operation's index. -
.gitignorefiles are now respected even if they aren't materialized in the
working copy because they are excluded by the sparse patterns. Previously,
ignored files could become tracked in a sparse working copy.
#2289 -
In-tree ignore files (
.gitignore) are no longer read through symlinks,
matchinggitbehavior. Such files are now silently skipped instead of having
their symlink target applied.$GIT_DIR/info/excludeandcore.excludesFile
are unaffected and still follow symlinks, asgitdoes.
#7161 -
jj workspace listtemplates are now labeled withworkspace name,
workspace root, etc.
Contributors
Thanks to the people who made this release happen!
- Aaron Bies (@slerpyyy)
- Austin Seipp (@thoughtpolice)
- Bartok9 (@Bartok9)
- Brice Figureau (@masterzen)
- Bryan O'Sullivan (@bos)
- Caleb White (@calebdw)
- David Rieber (@drieber)
- Farid Zakaria (@fzakaria)
- Gabriel Goller (@kaffarell)
- Gasper Stukelj (@mirkomartn)
- Jakub Stasiak (@jstasiak)
- JamBalaya56562 (@JamBalaya56562)
- Joseph Lou (@josephlou5)
- Karnajeet Gosavi (@kg290)
- LOG (@logarithmone1128)
- Martin von Zweigbergk (@martinvonz)
- Matt Stark (@matts1)
- Mustafa Officewala (@genericusername2709)
- pederbe (@pederbe)
- Philip Metzger (@PhilipMetzger)
- Pro (@twistedfall)
- Remo Senekowitsch (@senekor)
- Sami Hiltunen (@SamiHiltunen)
- sofia (@badp)
- Stephen Jennings (@jennings)
- Vincent Ging Ho Yim (@cenviity)
- xtqqczze (@xtqqczze)
- Yannik Sander (@ysndr)
- Yuya Nishihara (@yuja)
- Jujutsu can now colocate workspaces besides the default one by creating Git
-
🔗 Simon Willison Claude Haiku 5.5 rss
As previously promised, here's Anthropic's new fast, low cost model: Introducing Claude Haiku 5.5.
The previous Haiku, 4.5, was very much showing its age. It came out almost a year ago, and was priced at $1/million input and $5/million output - relatively expensive even back then, and a full 10x the price of OpenAI's GPT-6 Luna, released last month.
The new Haiku exactly matches the price of GPT-6 Luna - $0.10/$0.50 - up to 100,000 tokens. Beyond 100,000 tokens the price increases 5x to $0.50/$2.50. Luna itself has a price increase at 272,000 tokens but only to $0.20/$0.75.
Haiku 5.5 also uses a new, less generous tokenizer. My Claude Token Counter tool shows that the same long prompt uses around 1.25x as many tokens with Haiku 5.5 compared to Haiku 4.5, so there's a hidden price increase there.
If your workloads fit in 100,000 tokens, Haiku is the same price as Luna and reports higher benchmark scores. Above 100,000 tokens, Luna looks like a much better deal.
The most recent release of llm-anthropic finally fixed it so I don't need to ship a new version of that plugin for every new model. I tested the new model like this:
llm install -U llm-anthropic llm anthropic refresh llm -m claude-haiku-5.5 "Generate an SVG of a pelican riding a bicycle" -o thinking_effort lowPelicans
Here are pelicans for low, medium, high, xhigh, and max. The new Haiku doesn't let you disable reasoning, and defaults to
medium. I got a good bicycle frame for everything beyondlow. The low effort pelican cost 0.0936 cents and took 7 seconds.This
maxeffort pelican (with a reasoning trace that starts "This is the classic pelican-on-bicycle SVG test...") took 5 minutes 9 seconds to generate, but still only cost me 3.3826 cents:
(Since the reasoning trace exhibits awareness of the benchmark, here's Generate an SVG of an armadillo in fishnet tights jaywalking on Mars (on xhigh), and the same prompt against some other recent models. Background on that.)
For comparison, here's the pelican I got a year ago from Haiku 4.5 (for 0.7583 cents - Haiku 4.5 did not support reasoning levels). It sucked at drawing pelicans:

And a generous API credit scheme for subscribers
In addition to Haiku 5.5, Anthropic announced today that they are halving the price of cache reads for Sonnet 5.5. They've also added API credits to subscription plans:
Second, this week, we’ll roll out a new monthly API credit to all Max and Team subscribers for use on the Claude Platform. Max 5x users will get $100 in credits per month, Max 20x users will get $200, and Team subscribers will receive up to $500, pooled across their users.
Claiming this is pleasantly easy: navigate to Settings -> Billing and select the API organization that should benefit from the credits every month:

The API credits exactly match the cost of the subscription itself. This is really generous - it makes it much easier for subscribers to use the API. Anthropic also let you disable auto-reload for the API, with the consequence that "API requests will stop when your balance runs out" - exactly what you want if you're planning to burn through those API credits without risk of a nasty billing surprise.
Note that the monthly credits do not roll over - use them or lose them.
OpenAI still allow you to use your Codex subscription for personal API use, which works out as a better deal for heavy API users. This new credit scheme goes at least some way to overcoming that difference.
You are only seeing the long-form articles from my blog. Subscribe to /atom/everything/ to get all of my posts, or take a look at my other subscription options.
-
🔗 r/Harrogate New to Harrogate and wondering about traffic around the centre rss
Due to getting a new job in Leeds and moving out of a relatively leafy part of London, I’m thinking of a move to Harrogate and looking at places not too far from the town centre. I’ve been looking at places on Ripon Road (and the roads just off it) near(ish) to the Doubletree Hilton. Every time I look at Google Maps, it seems to show heavy , often stationary, traffic in both directions and particularly around Crescent Road. Would living around there be an eternal traffic nightmare, or just at peak times of the working week?
submitted by /u/Creepy_Ad6865
[link] [comments] -
🔗 exe.dev Crossing the Hyper-Thread Boundary rss
Modern processors have multiple CPU cores, and the physical CPU cores in turn have two logical CPU cores. The latter permit the physical core to run two independent programs simultaneously, an approach known as hyper- threading. Hyper-threading uses CPU resources more efficiently, but it exposes transient execution vulnerabilities, in which a program running on one hyper-thread is able to extract data from the program running on the other hyper-thread. An example of such an attack is MDS.
exe.dev runs virtual machines on behalf of different users. We need to protect against the possibility of one user exploiting a vulnerability to extract data from another user running on a different hyper- thread of the same CPU core. Fortunately the Linux kernel supports core scheduling cookies to control which processes are permitted to share a CPU core.
It is straightforward to give each VM an independent cookie, meaning that two different VMs never run on the same CPU core. However, in practice this leads to measurably inefficient use of physical CPU cores. So we instead implemented a more efficient, but still secure, mechanism: the VMs of each team use an independent cookie. This means that two different VMs from a different team (or from a different user for users not on teams) never run on the same CPU core. To put it another way, we assume that different members of the same team trust each other.
Of course, for various reasons, some teams do not have trust among all their VMs. An admin of those teams can run
ssh exe.dev team settings core-sharing offto prevent their VMs from sharing CPU cores. (Currently VMs that do not belong to a team never share cores with other VMs, even VMs from the same user; if that changes someday we will introduce a similar setting for individuals.) -
🔗 backnotprop/plannotator v0.28.7 release
Follow @plannotator on X for updates
Missed recent releases? Release | Highlights
---|---
v0.28.6 | Ask AI names the lines you selected, reorder quick labels and edit their emoji, Pi fixed-port crash fixed, OpenCode 2 subagent notice fixed
v0.28.5 | Several files in one review, theplannotatortool on Pi and OpenCode 2, decisions name the exact file, Ask this session reconnects after sleep
v0.28.4 | Comments come back after the agent edits a file, the agent can list and close its reviews, PR-description feedback no longer dropped, Pi thinking levels
v0.28.3 | Typing into your agent while it answers an Ask no longer streams into Plannotator
v0.28.2 | Question cards show wrapped choices and tables, pinned images named by their file, Done with nothing to send starts no agent turn
v0.28.1 | Ask AI from diagram comments, pinned images named for the agent, OpenCode 1 URL toasts and one reply per feedback, folder feedback sent once
v0.28.0 | Ask this session in Claude Code, Pi and OpenCode 2, Claude Code mod on by default, Pi plan review no longer blocks, first-run demo
v0.27.25 | Code review works withcolor.diff = always, Bitbucket review fixes, wide tables no longer collapse in Firefox, install script fix
v0.27.24 | Image previews stay in the all-files view, PR comment previews open on the commented line
v0.27.23 | Bitbucket Cloud PR review, Question UI for answering agents in place, opt-in auto-update, viewed files remembered, review another repo or worktree
v0.27.22 | Plans open the docs they link to, Claude review jobs locked down, Code Tour on Linux without Claude's sandbox, Pi reviews the latest planWhat's New in v0.28.7
Seven PRs, two from community contributors, one of them a first-time contributor. The main fix stops Ask this session from handing your unsubmitted comments to the agent as instructions. Other changes: comments on images survive a reload, other plugins' commands no longer carry Plannotator's name, and every Plannotator server now checks the hostname it was reached by.
Ask this session keeps your drafts as drafts
When you asked a question through "Ask this session", Plannotator also sent the comments you had not submitted yet, in the same format as submitted feedback. The agent read them as instructions and could act on them before you pressed Submit. In the report, it created GitHub issues from draft comments. On Submit, the same comments then arrived a second time.
Draft comments are now sent once, as a short list clearly marked as drafts you have not submitted, with an instruction not to act on them. Your question always comes last. The list goes again only when it changes. Submit delivers your feedback once, as before. This applies on Claude Code, Pi and OpenCode, and to the annotate agent terminal. The separate Ask AI providers are unchanged. (#1749, closing #1748, reported by @krizman)
Comments on images come back after a reload
In HTML annotate, a pin on an image, video or embedded page that had no
iddisappeared after a reload or an HTML refresh, which made image comments in generated reports hard to keep. Such pins are now found again by the file they show. A pin follows its image when other images are added around it. If the image's source changes, the pin shows as no longer matching. If the same source appears twice, the pin is not saved, so it can never land on the wrong copy. Live-app sessions behave the same way. (#1747 by @pro- vi)Other plugins' commands no longer carry Plannotator's name
With the Plannotator mod on, Claude Code labelled the output of any other plugin's slash command as if Plannotator had helped answer it (for example
plannotator+frontend: …). The mod now listens only to its own commands./plannotator-review,/plannotator-annotateand/plannotator-lastwork as before, with or without the Plannotator skills installed. (#1741 by @rushelex, closing #1740)Every server checks the hostname it was reached by
Plannotator's review servers now refuse requests that arrive under an unexpected hostname. A normal local session accepts only
localhostand loopback addresses. Remote mode also accepts IP addresses, your configuredPLANNOTATOR_URL_HOSTand the machine's own name (includingname.local).--tailscalesessions accept their tailnet name, and code-server and Coder proxies are recognised fromVSCODE_PROXY_URI. VS Code, SSH and Docker port forwarding, Codespaces and dev tunnels keep working as before.If you reach a remote-mode session by another DNS name (a server's public name, for example), you now get a 403 that says what to do: set
PLANNOTATOR_URL_HOST, or list the name in the newPLANNOTATOR_ALLOWED_HOSTS(comma-separated;*turns the check off). (#1742)Additional Changes
- OpenCode 2: extra words no longer stop annotate.
/plannotator-annotate . notes.md(orplease notes.md) opened nothing and showed nothing. It now opens the file, and a slash command that fails shows the reason in the session instead of only in OpenCode's log. On OpenCode 1 this applies when the plugin uses the CLI runtime; the default embedded runtime is unchanged (#1739). - Question cards for host apps.
@plannotator/ui0.52.0 lets an app embedding the question cards turn "Records a decision" on and off from the card's header and hide the card's own status tag. Plannotator's own cards are unchanged (#1744).
Install / Update
macOS / Linux:
curl -fsSL https://plannotator.ai/install.sh | bashWindows:
irm https://plannotator.ai/install.ps1 | iexClaude Code Plugin: The plugin and the
plannotatorbinary update separately, so run the install script above as well. In a terminal:claude plugin marketplace update plannotator claude plugin update plannotator@plannotatorThen restart Claude Code. Inside Claude Code, run
/plugin marketplace update plannotator, then open/plugin→ Installed → plannotator → Update now.Pi:
pi update --extensionsOpenCode: Re-run the install script above. It now also clears the OpenCode 2 plugin cache.
What's Changed
- Ask this session: unsubmitted annotations are never sent as feedback by @backnotprop in #1749
- fix(annotate): pins on images without an id restore by their source by @pro-vi in #1747
- fix(mod): match command.run by command name by @rushelex in #1741
- Validate the Host header on every request (defense in depth) by @backnotprop in #1742
- OpenCode 2: /plannotator-annotate with extra words opens the file, and failures are shown by @backnotprop in #1739
- ui: host-controlled decision toggle and status tag on question cards by @backnotprop in #1744
- Groundwork for a feature that is not released yet, hidden from help and the agent skill, by @backnotprop in #1745 and #1750
New Contributors
Contributors
@pro-vi found that comments on images without an id were lost on every reload, and wrote the fix that finds them again by their source. @rushelex reported the
+plannotatorlabel on other plugins' commands and fixed it with a one-line matcher. It is their second contribution, after the VS Code clipboard fix in #970.Community:
- @krizman reported the Ask this session draft problem with exact steps, the text the agent received, and the likely cause (#1748).
Full Changelog :
v0.28.6...v0.28.7 - OpenCode 2: extra words no longer stop annotate.
-
🔗 r/Harrogate Fresh Stop - what’s going on rss
Has anyone been to the shop near Asda? It’s one of those little grocery/corner-shop type places where you’d normally expect them to sell Pepsi, Coke, vapes, snacks, stuff like that.
I went in a few times now there was basically nothing on the shelves. It was really odd. Does anyone know what the deal is with it? Is it some kind of front/drug thing, or is there a completely normal explanation I’m missing?
submitted by /u/FrontRaspberry5060
[link] [comments] -
🔗 backnotprop/plannotator v0.28.6 release
Follow @plannotator on X for updates
Missed recent releases? Release | Highlights
---|---
v0.28.5 | Several files in one review, theplannotatortool on Pi and OpenCode 2, decisions name the exact file, Ask this session reconnects after sleep
v0.28.4 | Comments come back after the agent edits a file, the agent can list and close its reviews, PR-description feedback no longer dropped, Pi thinking levels
v0.28.3 | Typing into your agent while it answers an Ask no longer streams into Plannotator
v0.28.2 | Question cards show wrapped choices and tables, pinned images named by their file, Done with nothing to send starts no agent turn
v0.28.1 | Ask AI from diagram comments, pinned images named for the agent, OpenCode 1 URL toasts and one reply per feedback, folder feedback sent once
v0.28.0 | Ask this session in Claude Code, Pi and OpenCode 2, Claude Code mod on by default, Pi plan review no longer blocks, first-run demo
v0.27.25 | Code review works withcolor.diff = always, Bitbucket review fixes, wide tables no longer collapse in Firefox, install script fix
v0.27.24 | Image previews stay in the all-files view, PR comment previews open on the commented line
v0.27.23 | Bitbucket Cloud PR review, Question UI for answering agents in place, opt-in auto-update, viewed files remembered, review another repo or worktree
v0.27.22 | Plans open the docs they link to, Claude review jobs locked down, Code Tour on Linux without Claude's sandbox, Pi reviews the latest plan
v0.27.21 | Remote and phone sessions load several times faster, real Request changes on GitHub, model pickers show real names, OpenCode fixesWhat's New in v0.28.6
Six PRs, two of them asked for by the community. This is a fix release: Ask AI tells the agent which lines you selected, quick labels can be reordered and given a new emoji, and a few problems found while testing 0.28.5 on Pi and OpenCode are fixed.
Ask AI says which lines you selected
When you select text in a markdown document and ask a question, the question now names the lines, for example
Source: /path/to/doc.md, line 41, orlines 41–44when the selection spans several lines or blocks. Before, it carried only the file path, so when the same phrase appeared twice the agent could not tell which one you meant. The line numbers are the same ones the exported comment for that selection prints, so the question and your feedback point at the same place. This works in plan review and annotate, including linked documents and folder sessions, and through the side chat, Ask this session and the annotate agent terminal. (#1732, closing #1731, requested by @de-tre)Reorder quick labels and change their emoji
Settings → Labels now has Move up and Move down buttons on each quick label, and the emoji is an editable field. The list order decides which Alt/⌥ number applies a label, and the key hint on every row updates as you move things. The emoji field takes exactly one emoji (flags, skin tones and combined emoji included); anything else is shown as invalid and never saved. Two labels with the same text no longer get mixed up when you edit one of them. Your saved labels keep the same format, so nothing needs migrating. (#1738, closing #1736, requested by @RobertoArtiles)
Pi no longer crashes when a fixed port is busy
With
PLANNOTATOR_PORTset in a local Pi session, opening a second/plannotator-annotatewhile one was already open killed Pi withEADDRINUSE. The new review now takes the port over from the old one, as remote mode already did, and Pi stays up. (#1733)OpenCode 2: a review opened by a background subagent no longer adds a
stray turn
With the
plannotatortool turned on, a background subagent that opened a review posted its "session ready" link into the main session while it was idle. The next time the main session woke up, the model answered that link line as its own turn, and the exchange stayed in every later request. The link now goes to the session that called the tool, while that call is still open, so it never becomes a turn of its own. Decisions were always delivered correctly and still go to the main session. (#1734)Additional Changes
- Review names match what opened. On Claude Code,
/plannotator-annotate . a.mdopens onlya.md, but the mod named the review "2 files: ., a.md" in its reply, status line and decision. Reviews are now named from what the CLI actually opened, on Claude Code and for OpenCode 2's tool (#1737). - Contributors' OpenCode setup left alone. Running
bun installin a Plannotator checkout overwrote the developer's own OpenCode commands and skill. The OpenCode plugin's install step now copies them only when the package is installed undernode_modules, and it works on Windows. SetPLANNOTATOR_OPENCODE_POSTINSTALL=1or0to force it either way (#1735).
Install / Update
macOS / Linux:
curl -fsSL https://plannotator.ai/install.sh | bashWindows:
irm https://plannotator.ai/install.ps1 | iexClaude Code Plugin: The plugin and the
plannotatorbinary update separately, so run the install script above as well. In a terminal:claude plugin marketplace update plannotator claude plugin update plannotator@plannotatorThen restart Claude Code. Inside Claude Code, run
/plugin marketplace update plannotator, then open/plugin→ Installed → plannotator → Update now.Pi:
pi update --extensionsOpenCode: Re-run the install script above. It now also clears the OpenCode 2 plugin cache.
What's Changed
- Ask AI: name a text selection's source lines in the question by @backnotprop in #1732
- Settings → Labels: reorder quick labels and edit their emoji by @backnotprop in #1738
- fix(pi): attach the annotate agent terminal after listen so a busy fixed port cannot crash Pi by @backnotprop in #1733
- fix(opencode): a tool launch's session-URL notice never wakes an idle root by @backnotprop in #1734
- fix(mod, opencode): name a review by what the CLI opened, not the typed words by @backnotprop in #1737
- fix(opencode): skip the plugin postinstall inside the monorepo (cross-platform) by @backnotprop in #1735
Community
- @de-tre asked for the selected lines to be included in Ask AI questions, so the agent knows which occurrence of a phrase they meant (#1731).
- @RobertoArtiles asked for a way to reorder quick labels and change their emoji (#1736).
Full Changelog :
v0.28.5...v0.28.6 - Review names match what opened. On Claude Code,
-
🔗 matklad On Git Refs rss
On Git Refs
Oct 7, 2026
I have recently improved my mental model of Git. Consider these two git commands:
$ git fetch origin master $ git switch -c my-feature origin/masterDo you understand why is it
origin masterin one command, andorigin/masterin the other? I didn’t, until a few weeks ago!My understanding was that git is a content-addressable database. Git stores commits, a commit is identified by the hash of its content, and the content of a commit is, primarily:
- a memory-less snapshot of a state of the codebase at a given point in time,
- a list of (hashes of) parent commits.
That was enough git for me to understand
git logoutput and get me out of any botched rebase without having to re-clone the repo (For roughly half of my career, I was re-cloning the repo. No shame in that! Learning git is useful, but it’s not the highest priority thing to learn when you start).
I now understand that git not only comes with an append-only (“immutable”) content-addressable database, but is also a boring mutable key-value store.
Git has a mutable map whose keys are strings, and whose values are content- addressed objects. The keys are conventionally formatted as file system paths, and you can usually inspect the state of the mapping by listing
.git/refsdirectory:$ eza -T .git/refs .git/refs ├── heads │ ├── make │ ├── master │ ├── my-feature │ └── pbd-adt ├── origin ├── remotes │ └── origin │ ├── context-switches │ ├── gh-pages │ ├── HEAD │ ├── make │ └── master └── tags $ cat .git/refs/heads/master b59148228e52f7c615ead7fdd4e91001994ad50f $ git show-ref refs/heads/master b59148228e52f7c615ead7fdd4e91001994ad50f refs/heads/masterWhat makes this
refsKV infrastructure confusing is that:- It powers many distinct user-visible git features, but refs themselves are an implementation detail.
- It is a bit of a leaky abstraction, refs are almost invisible in the day-to-day usage.
- Git CLI uses shorthand notation for refs and many default arguments, which makes it not obvious that a particular CLI argument is a ref.
- And, as usual, git likes to give several names to one thing, and re-uses the same name for distinct things.
Branches, tags, and git notes are all just refs!
The structure becomes much more obvious once you elaborate all CLI shortcuts. The original command
$ git fetch origin masterthen becomes
$ git fetch \ https://github.com/matklad/matklad.github.io \ refs/heads/master:refs/remotes/origin/masterThe first argument of
fetch(https://...) is a location of a remote repository. Git will “dial” that address, and will transfer some data from that computer locally over the network.The second argument is a
source:targetpair of string keys (refs). Thesourceis a key on the remote repo, thetargetis the name of a local key, and fetch as a whole asks git to read a value from a remote repository and save it locally under a different name.To avoid typing repository URLs all the time, git assigns them symbolic names, with
originbeing the conventional name for the primary remote repository:$ git fetch origin \ refs/heads/master:refs/remotes/origin/masterrefs/heads/masteris a fully elaborated name of a branch on the remote repo. That is, branchmy-featureis just arefs/heads/my-featureref. It could have beenrefs/branch/my-feature, but it isn’t :)I don’t know the specific shorthand rules, but, generally, git allows you to spell only the suffix of a ref:
$ git fetch origin \ master:refs/remotes/origin/masterrefs/remotes/origin/masteris the name of the local ref we’ll use to store the result. It would seem natural to just use the same name locally as the one on the remote, but this only works if there’s a single remote. If there are two upstream repositories (for example, your fork, and the original repo you forked from), their ref names will collide. That’s why we want to namespace the refs for remote calledfoounderrefs/remotes/foo. Andoriginis just a conventional name for the remote in simple setups.Again, it would be more natural to directly mirror remote ref structure locally:
refs/ heads/my-branch -> refs/ remotes/origin/ heads/my-branchbut git strips the redundant heads component. And this
-heads,+remotes/$remotemapping is built in, which compresses the command to$ git fetch origin masterIt’s worth reflecting why it works this way. Git model is offline first. What’s more, it assumes explicit synchronization points. Rather than synchronizing with the remote repository in background when there’s connectivity, git requires explicit
fetchandpushoperations to transfer bytes over the wire. In this paradigm, it is useful to model the state of the remote party at the moment when we talked to them the last time. Theory of mind!This hopefully deconfuses git’s concept of local and remote branches. Consider the
mainbranch. It exists on the remote namedoriginasrefs/heads/main. When you synchronize your local repository withorigin, you getrefs/remotes/origin/main— you current best knowledge about the the state ofmainon theorigin.And then there’s your local
refs/heads/main. It typically starts pointing at the same commit asrefs/remotes/origin/main. But, when you make a commit,refs/heads/mainadvances, butrefs/remotes/origin/mainstays the same.When you try to push your local commit to origin, you will get a conflict, if the
mainbranch on theoriginadvanced in the meanwhile. In that case, git automatically updatesrefs/remotes/origin/main(as that’s just a local mirror of the remote state), but then it’s on you to updaterefs/heads/mainand push it again.Revisiting the full example:
$ git fetch origin master $ git switch -c my-feature origin/masterThe first command looks up the URL for the
originremote in.git/configand makes a network request to that machine. As a result, the localrefs/remotes/origin/mastergets updated to the same commit asrefs/heads/masterremotely (the commit and its ancestors are transferred locally as a result).The second command creates a
refs/heads/my-featureref (a branch), whose starting point isrefs/remotes/origin/master. It is an example of a leaky abstraction.The second argument there is a (shorthand of a) ref, so you can do
$ git switch -c my-feature \ refs/remotes/origin/masterBut, although the first argument creates a ref, it isn’t a ref itself. In other words, if you try to elaborate it as well
$ git switch -c refs/heads/my-feature \ refs/remotes/origin/masteryou’ll get
refs/heads/refs/heads/my-featureThat’s all! I am pretty sure this isn’t particularly useful, but maybe it is interesting!
-
🔗 New Music Releases Philip Glass - Philip Glass: Music for Film rss
Philip Glass - a new release is available:
- 2026-10-07: Philip Glass: Music for Film (Album)
Amazon: Canada | Deutschland | France | United Kingdom | United States
Visit muspy for more information.
-
- October 06, 2026
-
🔗 IDA Plugin Updates IDA Plugin Updates on 2026-10-06 rss
IDA Plugin Updates on 2026-10-06
Activity:
- capa
- disrobe
- f6fae606: test(dotnet): consolidate integration target
- d52b6b70: test(py-decompile): consolidate integration target
- 7743c9f6: test(jvm): consolidate integration target
- 1f3c2e06: fix(xtask): simplify consolidated audit paths
- 9b22833a: fix(xtask): match consolidated module filters
- 8cf5b4e9: fix(xtask): scope explicit test module attributes
- d542420b: fix(xtask): audit consolidated test modules
- b1767a0c: fix(ci): quote consolidated test filters
- 42abf9ec: test(cli): consolidate integration target
- 5e978218: test(native): consolidate integration target
- e7d81657: chore(xtask): regenerate the published figures
- cc3e426e: test(js-deob): consolidate integration target
- 13ee9ce6: test(testkit): load bounded runtime fixtures
- distro
- 41125aba: Build 14 PPA packages for arm64
- mcrit-plugin
- rhabdomancer
- 3a4dff38: doc: improve readme and changelog before release
- 58e9a077: doc: futher improvements
- 04dacdf1: doc: minor improvements
- 8910d093: feat: label call locations in library code recognized by IDA with `(l…
- a1f8bb58: feat: label call locations in library code recognized by IDA with `(l…
- 2d778674: doc: add todo item for following calls through thunks outside
.plt… - cd7ad31c: fix: stop marking instructions that fall through into a bad function,…
- 75545f8b: fix: match the glibc aliases that IDA may pick over the plain names i…
- f6b383e6: refactor: stub concept unification
- 1c4146c4: doc: add a note about fortified function handling
- fc1b1865: fix: match import stubs that IDA names with a numeric suffix and Univ…
- Signature-Locator
- 01539af9: feat: support manual input without length restriction and multi-langu…
- symbolicator
- c2ed4ef5: chore: refresh kernel signatures with IDA 9. + latest 27.2 beta KDK
-
🔗 Evan Schwartz Scour - September Update rss
Hi friends,
In September, Scour scoured 1.2 million articles (up from ~880,000 in August) from 28,568 feeds. Also, welcome to the 116 new users who signed up since my last product update email!
Here's what's new in the product:
📚 Library and Article Tabs
You can now find all of your saved, loved, and liked posts, as well as your full reading history, in the Library section.
Also, if you click Read on Scour for any article, that page now has tabs for the article's content, other posts that it cites and that cite it, and the feeds it was found in. Here's an example for a widely cited post.
🎓 Expertise Level
Scour now tries to determine the level of expertise each post assumes and infers the level of expertise you have per topic (based on the wording of your interest is and the types of articles you click on or like). At least for me, this means I'm seeing far fewer beginner Rust questions from Reddit showing up in my feed. (For those in tech, this is powered by Jev.)
🔎 Search for Feeds and People
Scour's Search will now show you results for feeds and authors, in addition to posts that match your query.
Relatedly, you can now follow individual authors as sources and Scour will try to show you their posts from any website they publish on.
🗑️ Detecting More Junk
Scour now detects and hides more junk, ranging from sales pages and SEO garbage to uninformative link roundups and low-value discussion threads. You should see more high-quality content in your feeds. By my current count, about 1 in 10 posts being shown before was some kind of junk that Scour now hides.
⚡ Faster Feed
I continue to obsess over making Scour feel super fast and snappy. In September the slowest feed loads got about 7x faster (p99 went from 2.1 seconds to 282 milliseconds) and the median feed load time got 2x faster (p50 went from 90 ms to 40 ms). This is also while ranking about 1.4x as much content as the month before.
🪦 RIP Reddit Feeds
Unfortunately Reddit announced their plan to turn off RSS feeds on November 13th. This is how Scour checks which discussions are happening on Reddit and finds articles posted on different subreddits. After November 13th you'll no longer see links to the Reddit discussions from Scour 😢.
🔖 Some of My Favorite Posts
Here are some of my favorite articles I found on Scour in September:
- The biggest news in the tech / AI world was the release of TypeSafe's Jev model. These were some of the related articles that I found interesting:
- The Latent Space interview with TypeSafe's CEO, Jev: System One models for Prod, not God — with Diogo Almeida, CEO, TypeSafe AI.
- Fingerprints of Jev and Jev's Architecture Unmasked were interesting black box investigations into Jev's base model using the tokenizer and other externally visible properties.
- Sixteen Models Walk Into a Storefront gives a very nice breakdown of techniques that can be used to manipulate LLMs' assessments of which products to buy and how much to pay for products.
- Logo Design Trends in 2027 Favor Marks Someone Can Prove They Made. In the age of AI-generated glossy slop, this is no surprise, but it's a nice write-up.
- The engineering behind the US Strategic Petroleum Reserve. Quite random but an interesting read.
- I Judged My Mum for Overusing AI, Until I Caught Myself Doing Worse. I continue to appreciate Sid's commentary on the age of AI. This is very relatable.
Happy Scouring! - Evan
- The biggest news in the tech / AI world was the release of TypeSafe's Jev model. These were some of the related articles that I found interesting:
-
🔗 roboflow/supervision supervision-0.30.8 release
v0.30.8 — Sharper video, cleaner labels
Video, YOLO labels, masks, VLM parsing and mAP all get more accurate.
VideoSinkkeeps OpenCV's video quality when OpenCV isn't installed.from_yoloreads pose labels as boxes instead of polygons.MeanAveragePrecisionscores class-agnostic runs right when only one side has class IDs.from_vlmreturns one Florence-2 detection per object, not one per polygon.from_inferencemasks no longer drift up to a pixel up and left.
Drop-in upgrade. Without OpenCV, videos get larger; YOLO labels with a negative width or height now raise
ValueError.✨ Spotlights / highlights
sv.VideoSinkandsv.process_videokeep quality without OpenCVThe PyAV fallback left the encoder bit rate unset, so
mp4vandMJPGfiles came out at under half of whatcv2.VideoWriterwrites. It now uses OpenCV's rate settings for every codec except H.264. Files get larger andvp09may encode more slowly;codec="avc1"keeps files small where an H.264 encoder is available. (#2661)import supervision as sv video_info = sv.VideoInfo.from_video_path("in.mp4") with sv.VideoSink("out.mp4", video_info) as sink: # OpenCV not installed for frame in sv.get_video_frames_generator("in.mp4"): sink.write_frame(frame) # before: mp4v written at under half OpenCV's bit rate, visibly softer # now: same bit rate OpenCV's writer usesYOLO pose labels load as boxes
A pose row is a box followed by keypoints.
from_yoloused to parse the whole row as a polygon, giving wrong boxes and masks nobody asked for. It now reads the box and skips the keypoints thatkpt_shapedeclares. (#2655)ds = sv.DetectionDataset.from_yolo( images_directory_path="pose/images", annotations_directory_path="pose/labels", data_yaml_path="pose/data.yaml", # kpt_shape: [17, 3] ) # before: polygon-parsed boxes and masks # now: one box per rowClass-agnostic mAP with one-sided class IDs
A perfect match scored zero when only one side carried class IDs, such as SAM proposals checked against labeled ground truth. With
class_agnostic=True, both sides now count as one class.One Florence-2 detection per object
(#2648)
Florence-2 returns a segmented object as a list of polygons, one per connected region. An object split in two used to come back as two detections; the polygons now merge into one mask with one box around all of them.
Roboflow masks sit on the right pixels
(#2649)
from_inferencetruncated sub-pixel polygon vertices, shifting each mask up and left by up to a pixel. Vertices are now rounded, the way the COCO, YOLO, LabelMe and Pascal VOC loaders already do.🔄 Migration guide
No migration required for this release.
📝 Notable changes
🌱 Changed
sv.DetectionDataset.from_yoloraisesValueErrornaming the annotation file when a label has a negative width or height; it used to load a box withx_minpastx_max, which madeDetections.areanegative and skewed IoU and NMS.as_yolonow orders the corners of a reversed box before measuring, so it no longer writes a file the loader refuses. (#2663)
🔧 Fixed
sv.VideoSinkandsv.process_videowritemp4v,MJPGand other non-H.264 video at OpenCV's bit rate when OpenCV isn't installed. A frame rate of zero or less now raisesRuntimeErrorinsv.VideoSink, as it does with OpenCV. (#2661)sv.DetectionDataset.from_yoloreads the box of Ultralytics pose labels and skips their keypoints; akpt_shapeother than[K, 2]or[K, 3]raisesValueError. (#2655)sv.DetectionDataset.from_yolonames a malformed annotation line, one with too few values or, for OBB, not nine, in aValueErrorinstead of failing on an array shape. (#2665)sv.metrics.MeanAveragePrecision(class_agnostic=True)treats detections without class IDs as the same class as labeled ones, and unsigned class ID arrays no longer raiseOverflowErroron NumPy 2. (#2650)sv.Detections.from_vlmwithsv.VLM.FLORENCE_2merges the polygons of one instance into one detection for<REFERRING_EXPRESSION_SEGMENTATION>and<REGION_TO_SEGMENTATION>, and skips instances with no usable polygon. (#2648)sv.Detections.from_vlmwithsv.VLM.QWEN_2_5_VLorsv.VLM.QWEN_3_VLrecovers complete detections from a response cut off inside abbox_2darray or right after a complete object. (#2666)sv.Detections.from_inferencerounds polygon vertices to the nearest pixel before rasterising masks, and raisesValueErrorfor NaN or infinite vertices. (#2649)sv.xyxy_to_maskreturns an empty mask for a box entirely left of or above the image when its maximum coordinate is a negative fraction. (#2646)sv.Detections.get_anchors_coordinatescomputes axis-aligned midpoint anchors without integer overflow. (#2660)sv.LineZone.triggerages crossing history on frames whose detections lacktracker_id, so a reused track ID no longer creates a false crossing after the track expired. (#2644)sv.LineZoneAnnotator(text_orient_to_line=True)no longer raisesTypeErrorwithout OpenCV for lines drawn right to left. (#2659)- The Ultralytics, Inference and YOLO-NAS speed estimation examples measure elapsed time from frame indices; a vehicle missed in one frame of three was reported about 44% too fast. (#2654)
examples/speed_estimation/rfdetr_example.pyno longer raisesAttributeError:supervision._cv2now providesgetPerspectiveTransformandperspectiveTransform. (#2652)
🏆 Contributors
- Mohammad Hijjawi (@MohammadHijjawi97, LinkedIn) — fixed video quality without OpenCV, Florence-2 instance merging and Roboflow mask rounding.
- Kari Pikkarainen (@kari-pikkarainen, LinkedIn) — fixed speed estimation timing, the NumPy
flipfallback and the perspective-transform fallbacks. - Miral Amin (@aminmiral) — made YOLO loading reject negative extents and name malformed lines.
- NIKHIL (@Nikhi00718) — fixed
LineZonehistory expiry and anchor overflow. - kevin (@kevin9327) — made Qwen parsing recover from cut-off responses.
- A Aswanth Raj (@aswanth-07, LinkedIn) — fixed class-agnostic mAP.
- Devulapalli Naga Sri Vaishnavi (@Vaishnavi220506) — fixed masks for off-frame fractional boxes.
- JANG BYUNGKUN (@8rulerstar) — fixed YOLO pose label loading.
Automated contributions:@dependabot
Full changelog :
0.30.7...0.30.8 -
🔗 @HexRaysSA@infosec.exchange The upcoming IDA 9.5 adds 3️⃣ new decompilers and will deliver 🔟 platform mastodon
The upcoming IDA 9.5 adds 3️⃣ new decompilers and will deliver 🔟 platform updates.
The new decompilers:
◾ Android DEX
◾ Infineon TriCore
◾ Qualcomm Hexagon👉 Read the full blog to see the rest of the updates: https://hex- rays.com/blog/ida-9.5-three-new-decompilers
-
🔗 Hex-Rays Blog IDA 9.5: 3 new decompilers and 10 platform updates rss
Good tooling starts with solid fundamentals. When every instruction decodes correctly and every function reads as clean pseudocode, you can trust what IDA shows and put your time into the binary itself. That holds whether the one reading the output is a person or an agent driving IDA. IDA 9.5 brings that reliability to three new architectures and sharpens it on several familiar ones.

-
🔗 Andrew Ayer - Blog sourcespotter-authorize: Monitor Your Go Modules for Malicious Versions, Without the Noise rss
You can protect the users of your Go modules from supply chain attacks, such as a compromise of your GitHub account, by monitoring Go's checksum database (sumdb). Since the go command won't install a module unless its checksum is published in the sumdb, monitoring the sumdb lets you discover unauthorized versions of your modules. Since the sumdb is a transparency log, you can even detect if Google themselves go rogue and publish a malicious version of your module. Although you can only detect, not prevent, attacks, Go's Minimal Version Selection makes it possible to respond before most of your users have installed the malicious version. Dependency cooldowns, which are coming to Go, will make it even easier to respond in time.
How To Monitor
Source Spotter, which is operated by my company SSLMate as a free service to the Go community, provides Atom feeds listing all versions of your modules found in the sumdb. For example, this Atom feed returns all versions of modules under the src.agwa.name/ prefix:
https://feeds.api.sourcespotter.com/modules/versions.atom?module=src.agwa.name%2FYou can also get Prometheus-compatible metrics:
https://metrics.api.sourcespotter.com/modules?module=src.agwa.name%2FYou can scrape the metrics endpoint using Prometheus and alert in your usual way, subscribe to the Atom feed using your favorite feed reader, or use one of the free services that converts Atom feeds to emails.
Avoiding Noise
Discovery is only half the story. The more important half is deciding whether to alert on a discovered module. Approximately 100% of the records published in the sumdb are legitimate. If you're alerted every time a new version of your module is legitimately published, it will be very hard to notice the one time it's an attack.
I wasn't sure at first how Source Spotter could facilitate no-noise monitoring. Initially, I thought Source Spotter would need access to your module's Git repository so it could cross-check sumdb records against the repo's contents. But that seemed complicated, and it would fail to detect a compromise of your repository host. I also really wanted a solution that wouldn't require users to create accounts.
I finally found a lightweight solution that I really like. I'll show you how you use it, and then explain how it works.
First, you install a small command line tool called sourcespotter- authorize and generate a public/private key pair:
$ go install software.sslmate.com/src/sourcespotter/cmd/sourcespotter- authorize@latest sourcespotter-authorize -keygenRun sourcespotter-authorize again with the
-feed-foror-metrics-forflags to output Atom and Prometheus URLs for the module prefix you want to monitor (src.agwa.name/ in this example):$ sourcespotter-authorize -feed-for src.agwa.name/ https://feeds.api.sourcespotter.com/modules/versions.atom?module=src.agwa.name%2F&mldsa=efbcb2bcb2d4decdf1ad9cab3224b2dd0087b6c6877abe00448a00d31716dff1 $ sourcespotter-authorize -metrics-for src.agwa.name/ https://metrics.api.sourcespotter.com/modules?module=src.agwa.name%2F&mldsa=efbcb2bcb2d4decdf1ad9cab3224b2dd0087b6c6877abe00448a00d31716dff1These are the same URLs shown earlier, but with a new
mldsa=parameter in the query string which is the hash of the public key you generated in the previous step.Initially, these URLs return the same contents as the URLs without the mldsa= parameter. That's because you haven't marked any module versions as authorized yet.
To mark a module version as authorized, change into the Git repository for the module and run sourcespotter-authorize with a Git tag:
$ cd ~/src/snid $ sourcespotter-authorize v0.4.0Now, v0.4.0 of this module is omitted from the feed and metrics URLs.
The command accepts multiple tags as arguments, so you can authorize all the tags in a repo like this:
$ sourcespotter-authorize $(git tag)Moving forward, you should run sourcespotter-authorize any time you tag a new version. I've written a tiny shell script called gotag that runs git tag followed by sourcespotter-authorize:
#!/bin/sh -e git tag "$1" sourcespotter-authorize "$1"As long as you authorize every tag you create, the feed and metrics URLs will report zero unauthorized module versions, eliminating false positive alerts. A supply chain attacker can't hide their malicious versions from your feeds unless they also compromise sourcespotter-authorize's private key.
How It Works
sourcespotter-authorize -keygengenerates an ML-DSA-44 private key (stored under $XDG_CONFIG_HOME/sourcespotter-authorize). The SHA-256 hash of the corresponding public key goes in the mldsa query string parameter.When you authorize a tag, sourcespotter-authorize uses the golang.org/x/mod/zip and golang.org/x/mod/sumdb/dirhash packages to compute the checksums of the go.mod and module zip files for the tag. It formats the module path, version, and hashes for each tag as a go.sum file (no need to invent a new format here), and signs the file with your private ML-DSA key. Finally, it uploads a JSON object to the Source Spotter server containing your public key, the go.sum file, and the signature.
The Source Spotter server verifies the signature using the public key. If valid, it records the module path, version, and hash as being authorized by that public key, and excludes that module version from the feeds and metrics endpoints for the key.
The protocol is simple and documented if you want to implement your own client.
Why Not Just Sign Your Releases?
If you're generating a private key and signing stuff anyway, why not just sign your releases and distribute the signatures? After all, this is the traditional approach to stopping unauthorized software releases.
The problem with the traditional approach is that consumers of your software have to verify the signatures, which means they need to know what your public key is. Key distribution is a hard, hard problem, and most software ecosystems do a bad job at it. I would guess that in practice, few people actually verify signatures of software releases.
In contrast, the go command automatically verifies that every module it installs is listed in the checksum database, without the user needing to do anything. Although sourcespotter-authorize needs a private key, you only need to distribute the public key to Source Spotter (or whatever other monitor you might choose to use), rather than every single consumer of your software. That's a vastly simpler problem, especially when (not if) you need to rotate your keys.
Everyone who signs their software releases will need to rotate their keys soon, as quantum computers are expected to break RSA and elliptic curves within the next few years. An earlier version of sourcespotter-authorize used elliptic curve keys. After it gained ML-DSA support, upgrading my key was super easy - generate a new key, transfer my list of authorized versions to it, and finally update the URL in my feed reader:
$ sourcespotter-authorize -export > tmpfile $ rm ~/.config/sourcespotter-authorize/private_key $ sourcespotter-authorize -keygen $ sourcespotter-authorize -import < tmpfile $ sourcespotter-authorize -feed-for src.agwa.name/This was so much easier than communicating a new key to every consumer of my software would have been, and didn't require a lengthy transition period during which I was signing with two keys.
You can learn more about Source Spotter's module monitoring, get the source code on GitHub, or read my blog post about how Source Spotter also verifies Go's reproducible builds.
-
🔗 backnotprop/plannotator v0.28.5 release
Follow @plannotator on X for updates
Missed recent releases? Release | Highlights
---|---
v0.28.4 | Comments come back after the agent edits a file, the agent can list and close its reviews, PR-description feedback no longer dropped, Pi thinking levels
v0.28.3 | Typing into your agent while it answers an Ask no longer streams into Plannotator
v0.28.2 | Question cards show wrapped choices and tables, pinned images named by their file, Done with nothing to send starts no agent turn
v0.28.1 | Ask AI from diagram comments, pinned images named for the agent, OpenCode 1 URL toasts and one reply per feedback, folder feedback sent once
v0.28.0 | Ask this session in Claude Code, Pi and OpenCode 2, Claude Code mod on by default, Pi plan review no longer blocks, first-run demo
v0.27.25 | Code review works withcolor.diff = always, Bitbucket review fixes, wide tables no longer collapse in Firefox, install script fix
v0.27.24 | Image previews stay in the all-files view, PR comment previews open on the commented line
v0.27.23 | Bitbucket Cloud PR review, Question UI for answering agents in place, opt-in auto-update, viewed files remembered, review another repo or worktree
v0.27.22 | Plans open the docs they link to, Claude review jobs locked down, Code Tour on Linux without Claude's sandbox, Pi reviews the latest plan
v0.27.21 | Remote and phone sessions load several times faster, real Request changes on GitHub, model pickers show real names, OpenCode fixes
v0.27.20 | Mistral Vibe support, annotate gets the full Options menu and Settings, jj Commits panel, long lines wrap in plan code blocksWhat's New in v0.28.5
Twelve PRs, one from a first-time contributor. Agents on Pi and OpenCode 2 can now open Plannotator reviews themselves without holding the session, a review can cover several files at once, and every decision now names the exact file it is about.
Several files in one review
plannotator annotate spec.md ui/mock.html notes.mdnow opens one review of all the files, in the order you typed them. A header switcher shows "2 of 3", the Files tab keeps that order, and one decision covers the whole set, with the feedback split into a section per file. Ask AI knows which file you are on, and your unsent comments are kept per file. The same works when an agent passes a list of files to theplannotatortool. If one of the paths does not exist, nothing opens and the error names the missing file. Words that are not file paths keep their old meaning, soannotate look at notes.md pleasestill opensnotes.md. (#1718)The
plannotatortool on Pi and OpenCode 2, off until you turn it onOn Claude Code (with the Plannotator mod) agents already had a
plannotatortool. It is now on Pi and OpenCode 2 too. It matters when the agent opens Plannotator itself, for example when you ask it to "show me the HTML plan in Plannotator" or "open a Plannotator code review". Without the tool, the agent runs theplannotatorcommand in its shell, which blocks the session until you finish and leaves Ask AI unable to reach the session. With the tool, the review opens right away while the agent keeps working, your decision comes back as a message, Ask this session works from that review, and the agent can list and close the reviews it opened.The tool adds about 780 tokens to every request, so on Pi and OpenCode 2 it is off by default. The first Plannotator page you open there asks "Do you use Plannotator as a skill?" and turns it on for your next session if you say yes. You can change it any time with the new toggle in Settings → General (plan review, annotate and code review), the
agentToolkey in~/.plannotator/config.json, orPLANNOTATOR_AGENT_TOOL. On Claude Code the tool stays on by default, since Claude loads it only when it is needed; the same switch turns it off there without turning off the rest of the mod. Your/plannotator-*commands work the same either way. (#1714, #1715, #1724, #1725)The tool's description no longer includes the unfinished
replyaction, and Pi's bundled knowledge skill stays out of the model's context by default, keeping the promise from #842.Every decision names the file it is about
A user reviewing two different files that were both named
QUESTIONS.mdapproved one of them. The approval reached the agent as just "QUESTIONS.md — Approved." with no path, and the agent attached it to the other file. Every decision message now carries aTarget:line with the full path (or URL, PR, or folder), taken from the review that recorded the decision, and reviews that share a file name are labelled with their folder, such asreleases-2026-10-04/QUESTIONS.md. A code review decision names the PR that is on screen when you decide, even if you switched PRs inside the review.A related gap is closed too: when a review's port was reused (a fixed
PLANNOTATOR_PORT, remote mode, or rarely by chance), an old browser tab could submit a decision to the new review. Each review now has its own id, and a decision from a tab that belongs to a different review is refused with a "This review was replaced, reload" banner.plannotator sessionsnow lists each review's id and full path, and has a--jsonoption. (#1729)Ask this session reconnects after sleep
With the Claude Code mod, Ask AI in an open review could switch permanently to "This session is no longer available" after your computer slept or the network dropped for a few minutes. It now reconnects on its own, waiting a little longer between attempts while the review is unreachable. Running
claude --continuewhile the old window is still open no longer makes the two windows fight over the review: the window you are using takes it over, each decision is delivered exactly once, and plan approvals reach the right window. If you quit Claude Code after a decision arrived but before Claude saw it, you get a notice next time saying where it was saved. (#1727)Approve with a note in every gated review
Gated annotate reviews (
--gate, or the tool withgate: true) opened from Claude Code never offered "Approve with a note…", even though the note could be delivered. They do now, for markdown and HTML alike, on every host that delivers the note with the approval. Your note reaches the agent as an "approved with notes" message. A plain approval is unchanged: plain output is still exactlyThe user approved.One thing for scripts: when you approve with a note, plain output now prints the approved-with-notes message instead of that line. Scripts should use--gate --json. (#1728)Additional Changes
- Done says nothing was sent. Clicking Done with nothing to send showed "Feedback Sent". It now shows "Done, nothing was sent", and Close no longer claims a response was sent (#1730).
- Pi keeps the model you picked. Using
/treeto re-answer something during plan execution switched the model back to the plan's original one. It now keeps your choice; moving into or out of plan mode still applies the phase's model (#1723, fixes #1722). - Gruvbox diffs. Added lines in code review were grey on the Gruvbox theme. They are now Gruvbox green in light and dark (#1721).
- Safer settings. A web page in your browser can no longer change Plannotator's settings in the background, and saving settings from the VS Code panel works again, as do viewed-file progress and the Call Flow install there (#1724).
- Type checking now covers the Claude Code side's server code (#1720).
Install / Update
macOS / Linux:
curl -fsSL https://plannotator.ai/install.sh | bashWindows:
irm https://plannotator.ai/install.ps1 | iexClaude Code Plugin: The plugin and the
plannotatorbinary update separately, so run the install script above as well. In a terminal:claude plugin marketplace update plannotator claude plugin update plannotator@plannotatorThen restart Claude Code. Inside Claude Code, run
/plugin marketplace update plannotator, then open/plugin→ Installed → plannotator → Update now.Pi:
pi update --extensionsOpenCode: Re-run the install script above. It now also clears the OpenCode 2 plugin cache.
What's Changed
- feat(annotate): several files in one annotate review by @backnotprop in #1718
- feat(pi): the plannotator tool on Pi by @backnotprop in #1714
- feat(opencode): the plannotator tool on OpenCode 2 by @backnotprop in #1715
- feat(tool): agentTool switch with per-host defaults, Pi tool stability, drop the reserved reply action by @backnotprop in #1724
- feat(ui): agent tool switch in Settings and a one-time offer on Pi and OpenCode 2 by @backnotprop in #1725
- fix: decisions name their full target; refuse stale-tab decisions on a reused port by @backnotprop in #1729
- fix(mod): restart a dead Ask-this-session bridge; one watcher per launch across processes by @backnotprop in #1727
- fix(annotate): offer Approve with a note in every gated session that delivers it by @backnotprop in #1728
- fix(annotate): show a Done screen, not Feedback Sent, when Done sends nothing by @backnotprop in #1730
- fix(pi): keep the user's model on a /tree navigation within the same phase by @backnotprop in #1723
- Fix gruvbox positive diff colors by @TheEdgeOfRage in #1721
- fix(hook): type-check apps/hook/server and fix vibe-plan.ts narrowing by @backnotprop in #1720
New Contributors
- @TheEdgeOfRage made their first contribution in #1721
Contributors
@TheEdgeOfRage fixed the grey added lines in Gruvbox code review, with the color override the colorblind theme already uses.
Community:
- @jasonharrison reported the Pi model reset on
/treewith a clear repro on Oh My Pi (#1722).
Full Changelog :
v0.28.4...v0.28.5 -
🔗 Project Zero How to fix a bug in a fix rss
Project Zero often works with software vendors to remediate the vulnerabilities we report and provide broader guidance on making software more secure. Some vendors express concern about potential scenarios in which they are unable to fix vulnerabilities that are causing immediate user harm, due to limitations in their patch delivery systems. Since Project Zero encounters a wide array of systems designed to protect users in the case of exceptional exploitation scenarios, both through vendor discussions and security reviews, we want to share what we’ve learned.
This post provides an overview of systems in use by large vendors that allow them to remediate small volumes of vulnerabilities much faster than their typical update process. Our goal is to provide a reference for vendors seeking to implement or enhance the capabilities of such systems, and to encourage vendors to consider how they would fix an urgent vulnerability before they receive one.
Why patching takes time
Patching a vulnerability typically involves the following stages:
- Triage — a vulnerability report is received, validated, prioritized and assigned to a specific developer to be fixed
- Patch development — a software development team writes, reviews and commits code that fixes the vulnerability
- Testing — the patch is tested to ensure the vulnerability is remediated and the software still functions correctly when the patch is applied. This can include formal testing by a test team, automated testing and alpha and beta testing where a patch is shipped to a limited group of users for feedback on normal use.
- Partner review — some software updates require review by third parties before they can be shipped, due to relationships between the software vendor and other organizations, for example, carrier acceptance for some mobile updates.
- Delivery — the patch is delivered to and installed by end users
- Activation — sometimes an additional step, such as a system restart, is needed to switch the system to the updated software
Of course, this is a simplified picture. Patching can involve repeating steps, for example rewriting a patch if tests fail, or additional stages when third- party vendors are involved. However, this is a minimal set of steps most software updates require.
The challenges of emergency patches
While triage and patch development time contribute substantially to the speed at which vendors can generally patch vulnerabilities, they contribute less to emergency patch time. Triage is usually very fast in situations where vendors know they have an urgent problem, and patch development can be expedited based on priority. Only in rare circumstances, where a vulnerability is especially complex, or a vendor’s security team does not have a complete picture of their software’s components and who within their organization maintains them, have we seen urgent patches delayed in the triage or development phase. Likewise, partner agreements usually have exceptions for updates in emergency situations.
Most vendors’ patch speed is limited by the testing and delivery stages. Testing is important because all changes to software risk introducing unexpected behavior. The worst-case scenario is that inadequately tested software ‘bricks’ a device, causing it to malfunction in a way that it can no longer perform key functionality or receive software updates to remediate this. Buggy software updates have also led to situations where user data is corrupted or lost, and any decrease in software functionality after a security update makes users less likely to apply updates in the future.
The potential cost to vendors of shipping poorly tested updates varies depending on the nature of the underlying software. For example, if a mobile application is rendered unusable due to an update that corrupts local data or prevents it from launching, users can easily install the next version via an app store, and their data is usually saved on a remote server, so costs are limited to user support. Meanwhile, if a mobile device gets bricked, it needs to be returned to its manufacturer or place of purchase for repair, leading to substantial costs for the vendor and potentially the user.
The possibility of serious functional bugs is considered in the design of most patch delivery systems. Updates are often rolled out slowly, so that serious problems can be detected before they affect too many users. Often, patching vulnerabilities quickly and avoiding buggy patches are at odds with each other, requiring tradeoffs that prioritize one over the other.
A variety of other technical challenges can limit the speed of patch delivery. One is the design of the patching system. A common design is that devices probe for updates at a regular interval, leading to patch saturation being limited to that interval. ‘Push’ style update systems can deliver patches to all users faster, but generally require more infrastructure.
User behavior and environment can also be a barrier to patch propagation. Patches that require user interaction to install are often delayed by users, and network speed and data cost are also factors in installation rate. Updating many users at once, as opposed to over a period of time, can strain patch delivery infrastructure. Chrome and Microsoft have written about the challenges of updates requiring restart to install, as users are often reluctant to restart their system and restarts take time.
While testing delays and limitations of the patch delivery system affect all updates, the shorter time frame of emergency updates make them a larger contributor to the overall time it takes to deliver a patch.
Emergency patching methods
Feature flags
Feature flags are conditional statements in source with paths determined by values provided by a remote server. They are often used for A/B testing, but they can also be used for short term remediation of vulnerabilities in emergency situations. A widely publicized case of this was a serious 2019 FaceTime vulnerability, where Apple temporarily disabled Group Facetime with a feature flag. Several vendors have made at least some media codecs available in 0-click contexts controllable via feature flags, and can disable them in the case of active exploitation, falling back to another codec for realtime transmission.
The main benefit of feature flags as a vulnerability remediation method is that testing can be performed with each flag set in advance, so a fast update does not require shipping untested code. They can also be delivered to users much more quickly, as updating feature flags requires transmitting a very small amount of data.
Recently, Meta published a blog post on how they implemented a ‘dual stack’ library, in which two versions of the WebRTC video conferencing library were compiled into a single binary, with the version in use controllable via a feature flag. This technology enables rapid updates with less testing, as new versions can be shipped with the option to quickly move users back to the previous version if function problems occur. While Meta uses two versions of the same library, it would also be possible to create a ‘dual stack’ with two different libraries that implement the same features (for example, two H264 libraries), allowing an application to switch to a different library to render a specific vulnerability unreachable without loss of functionality in an emergency. This would require additional testing, but it is testing that can be performed up front. It could also be possible to have a second library that enables performance intensive mitigations that would block many possible bugs, such as ASAN, or enabling DCHECKs.
Filtering
Filtering is running a dynamically updatable ruleset, such as a regular expression, against untrusted input in order to block specific input that is required to reach a vulnerability. An example of this is Android’s Intent Firewall, which allows specific usages of an Android IPC mechanism called intents to be disabled based on rules in a dynamically updateable XML file, which enables blocking intents that can be used to exercise specific vulnerabilities. It was recently used to block vulnerabilities in third-party Android wallets.
Some platforms have endpoint detection software that can perform filtering on a wide variety of system input, for example Microsoft Defender on Windows systems, and Google Play Protect on Android devices. Rules that block specific exploits or make certain vulnerabilities unreachable can often be deployed to these applications very quickly. Endpoint detection requires parsing a great deal of untrusted input, often in privileged context, so these applications are not without risk, but in systems where they already exist, they are a potential method of emergency remediation.
As an approach, filtering is more flexible than feature flags. For feature flags to be effective, the vendor needs to determine what features they might want to disable in advance, and if this isn’t comprehensive, they might find themselves in a situation where a vulnerability can’t be remediated via feature flags. Meanwhile, filtering can be used to block a wide variety of inputs, even ones that have never been considered. The downside of filtering is that performing filtering frequently can decrease software performance, and at least some testing of new filters is required, and can’t be performed upfront without knowing the vulnerability that needs to be blocked, as it is possible to write filters that interfere with necessary system functions.
Alternate Channels
The network ‘channels’ used to deliver software updates to users can be slow for a variety of reasons discussed above. Vendors sometimes implement alternate channels that can be used to deliver smaller updates more quickly.
Android Pony Express (APEX) is an example of an alternate channel that can be used to ship updates to specific high-risk Android components faster than a full system update. It shortens the patch development time, as OEMs do not need to integrate updates to APEX components. APEX is available to OEMs, and can be used to update OEM- maintained libraries.
Several applications we’ve researched have the ability to update individual libraries outside regular updates, usually by having some flag that is regularly checked over the network, and then downloading the library and loading it with
dlopenor equivalent. While this is an effective way to avoid delivery-speed limitations of updates, it can also introduce critical vulnerabilities if libraries delivered in this way are not adequately verified by the client to have originated from the vendor. We encourage vendors to be cautious, and ensure that emergency update mechanisms of this variety have adequate security testing.Hotpatching
Some vendors have implemented update mechanisms that allow units of binary code smaller than libraries to be delivered and applied directly to the memory space of a running process. For example Linux supports Livepatch which enables kernel functions to be directly replaced in memory without a restart. Similarly, Windows’ hotpatch allows security updates that contain only updated functions to be delivered to users, and applied while the process is still running.
Hotpatching has the potential to deliver very flexible security patches to software very quickly, with no degradation of user experience, though it typically has some limits to the nature of patches it can deliver, for example, updates that require changing the definition of a structure shared between functions are sometimes not supported. Hotpatching has similar security downsides to alternate channels, and also carries the risk of introducing ways to bypass exploit mitigations, as it requires permissions to map pages with write-execute privileges at some point during patching. It also doesn’t address any of the testing challenges of rapid updates, just the delivery challenges.
The importance of emergency patching
LLMs are increasing the vulnerability discovery and exploitation capabilities of both attackers and defenders. A wider array of actors now have the ability to perform novel attacks at greater speed. In light of this, it is important for vendors to consider how to protect their users in the case of active exploitation. Rapid update mechanisms do not need to be heavyweight or be capable of fixing every possible bug and preserving perfect user experience in every scenario. Technologies like feature flags, filtering and alternate update mechanisms can remediate the most likely and severe vulnerabilities in the short term, while keeping devices reasonably functional for users.
It is urgent for vendors to plan how they will protect their users in the worst case scenario of widespread active exploitation. Actions taken now can greatly improve security outcomes for users. By taking stock of update mechanisms already available to them and implementing rapid remediation functionality where gaps exist, vendors can be better prepared for whatever the future holds.
-
🔗 syncthing/syncthing v2.1.6 release
Major changes in 2.1
-
Devices and folders can now be grouped in the GUI by setting the new
groupattribute. -
HTTP and HTTPS proxies with support for CONNECT can now be used, in
addition to the existing support for SOCKS proxies (the environment
variableall_proxy=https://...). -
Block indexing can be turned off for folders where it's more desirable to
optimise for reduced database size and overhead than minimal transfer
size (theblockIndexingattribute on folder configuration). -
GUI login session duration can be configured to be longer or shorter than
the default one week, or set to infinitely long. The cookie path can also
be adjusted. (ThesessionCookieDurationSandsessionCookiePath
attributes in the GUI configuration.)
This release is also available as:
-
APT repository: https://apt.syncthing.net/
-
Docker image:
docker.io/syncthing/syncthing:2.1.6orghcr.io/syncthing/syncthing:2.1.6
({docker,ghcr}.io/syncthing/syncthing:2to follow just the major version)
What's Changed
Fixes
- fix(model): introducers should not be able to add themselves to folders by @calmh in #10880
- fix(gui): add accessible labels to buttons in Edit Device modal (fixes #10873) by @tomasz1986 in #10874
- fix(model): properly error out when a temp file can't be created by @calmh in #10883
- fix: disable keepalive on most outgoing HTTP connections by @calmh in #10891
- fix(monitor): continue log writes when stdout is unavailable on detached Windows consoles (fixes #10882) by @Shablone in #10889
- fix(gui): improve header and button color contrast (fixes #10488) by @giri256 in #10815
- fix: use global short lived HTTP client with proper HTTP/2 support by @calmh in #10901
- fix: remove unnecessary delay in index transmission by @calmh in #10908
- fix(watchaggregator): properly convert floating-point time duration (fixes #10899) by @calmh in #10900
New Contributors
Full Changelog :
v2.1.5...v2.1.6 -
-
🔗 streamyfin/streamyfin v0.55.1 release
fix(ios): keep tab labels on their tabs when built with the iOS 27 SD…
-
🔗 HexRaysSA/plugin-repository commits sync repo: +1 release rss
sync repo: +1 release ## New releases - [ida-settings-editor](https://github.com/williballenthin/ida-settings): 1.3.1 -
🔗 Mitchell Hashimoto A Terminal Protocol for Program Status (OSC 7501) rss
(empty) -
🔗 Filip Filmar Razboj: a minimal GPU in TxHDL rss
Razboj is a minimal graphics rasteriser implemented in approximately one hundred lines of TxHDL. It reads a display list from memory and writes rendered pixels into a framebuffer over an AXI bus. TxHDL lowers the design to synthesizable Verilog and VHDL, and the build verifies both netlists against the software simulation trace. This post describes the rasteriser architecture and hardware design tradeoffs.
What it draws
The reference demonstration scene measures 64 by 64 pixels (4,096 pixels total). The scene consists of eight display list entries: a background clear, three rectangles, and four triangles. The rasteriser renders the entire scene in approximately 13,000 clock cycles. A verification harness reads the completed framebuffer from memory and saves the output image.
-
🔗 Armin Ronacher What is Codemode rss
More than a year ago I wrote a few posts here that recommended people not to load custom tools into their context (or MCP servers) but to just use more scripts. Most importantly I wrote that Code Is All You Need and I wrote about that MCP needs code. With Pi 1.0 we now added MCP support via Codemode which in some ways is a long time coming, but then also maybe somewhat surprising to some. So I want to share some updated thoughts on this blog on what this all means.
What Are Tools
When a harness like Pi provides tools for an LLM to call, it does so by supplying some tool definitions which then translate into some token structure on the server side. Whether a model is encouraged to call a tool is the result of the reinforcement learning process. Something I wrote about before if you want to learn more.
One of the reasons we strongly lean towards CLI and bash is because it allows easy composition of calls, and because the model also learns how the file system works when it's trained. So when it invokes a tool like
echo foo > /tmp/test.txtthe model also learns that after that tool call, there is now a file calledtest.txtin/tmp.However bash has one fundamental limitation which is that it can only compose programs that run. And there are some things, which are not programs, but native tools to the LLM and they sort of have to be.
The most obvious example here is
readorview_image. If a multimodal model needs to read an image, it cannot usecatfor that because the harness needs to inject the actual image payload into the protocol of the LLM.Another quite vivid example are sub agents. In order to spawn and orchestrate sub agents, it's tricky to avoid tools that are provided by the harness. While in theory the agent could provide a CLI tool that talks to the outer harness via environment variables and Unix sockets, it's a rather crude process. It however has another issue, and that is where the code runs.
Brains vs Hands
To better understand that, it's important to think a bit more about where all the bits and pieces run. There really usually are two different systems involved. The first is the brain, the harness: it runs on one machine. It's trusted. The second is often the same machine, but it's really where the tools are executing: the hands. In Pi we now call this the execution environment, but you can think of it as the target of all the operations.
Crucially what is important for us, is that there is a dividing line between the harness brain and the target environment that runs bash and executes the tools.
And splitting this in half has some really important consequences. For a start it means that they are running on different file systems and they have different levels of trust. If you for instance use a sandboxing solution like Gondolin your bash stuff will be sandboxed just fine, but the harness itself will not be.
Orchestrating The Harness
Which brings us to what Codemode really does: it's a way for the LLM to express and orchestrate complex operations on the harness side, but not the execution environment side. Codemode runs in the harness, in its own sandbox. In case of Pi it's running in QuickJS within a WASM runtime with intentional limitations: no network, no file system, no timers, limited RAM. The only way is to call more tools. You could also imagine that Codemode could run Scheme or some other language as well.
If you are not familiar with Codemode, it's basically just a way to issue tool calls from within some language, in our case JavaScript. That allows you to compose those calls without necessarily going through the LLM's context. Credit for naming goes to our friends at Cloudflare who coined it.
For instance if you issue a bash call as a regular tool call in the LLM, then we only throw the trailing 2000 lines into the context and if the agent wants more, it needs to look at the overflow file itself. If however the agent issues that invocation via Codemode, then the Codemode side gets larger outputs sent structurally.
Most importantly, because Codemode is JavaScript the agent can express concurrent operations and basic workflows. A common way in which you see agents now use this, is to first probe at 5-10 items from some tool response to see what it looks like, and to then write a Codemode script that processes the next n items.
Codemode also allows you to throw state into the transcript! That means that one Codemode invocation can stash away data, that the next call in the session can load again. And remember: this is on the harness host, not the sandbox.
In case of Pi, Codemode also allows you to issue calls that naturally do not make any sense in Pi's traditional interface. For instance if you want to generate images with an image model or you want to classify some text with a one shot classifier model, those Pi APIs are exposed via Codemode, but not via regular tools where they would just waste context.
What It Looks Like
So now that we talked a bunch about it, it's probably worth being a bit more explicit about it. Let's walk ourselves through some invocations of Codemode of recent Pi sessions of mine. Note that none of this code is human written. It's from real sessions of Pi, just re-indented for your viewing pleasure. The agent starts using Codemode automatically either because it's a task where the model already naturally picks up that tool, or because a user asked it to.
Note that Codemode is by default only enabled in Pi when MCP is enabled, but you can turn it on with
"defaultTools": ["+codemode"]in the settings. Just ask Pi to enable it for you.Generating Images
Let's start simple with image generation. Image generation is a feature that Pi supports in the AI SDK core, but it's not a tool that the agent can use. In the past the only way to use image models has been to write a bespoke extension or to have the agent run node itself and use the internal image APIs. However because we expose quite a few of the internal model APIs within Codemode, it means that the agent can use it:
const [painter] = await models.getAvailableOfType("image"); const result = await models.generateImages(painter, { input: [{ type: "text", text: "A cute little puppy sitting on a grassy " + "lawn, soft natural light, photorealistic" }], }); if (result.stopReason !== "stop") return result.errorMessage; for (const block of result.output) { if (block.type === "image") image(block); else text(block.text); }Note that the call to
image()sends the image back as image content to the LLM. On the harness side it feeds it directly into both the agent, as well as onto disk as a temporary artifact in case the agent wants to be able to pass that image back to bash.Classifying Things
Similar things apply to classifier models such as Jev. They also do not fit well into the workflows of an agent through the typical tools. But rather than making a bespoke tool available, Codemode just allows the agent to reach into the AI SDK and invoke those directly. Here you can see how Jev is used to mass process GitHub issues for a quick sentiment analysis:
const jev = await models.getModelOfType("classifier", "typesafe", "jev-latest"); const r = await tools.bash({ command: "gh issue list --state open --limit 100 " + "--json number,title,body,comments", }); const issues = JSON.parse(r.output); const results = await Promise.all(issues.map(async (issue) => { const res = await models.classify(jev, { state: { title: issue.title, body: (issue.body || "").slice(0, 4000), comments: issue.comments.slice(-5).map(c => c.body.slice(0, 800)), }, questions: { sentiment: { type: "choice", instructions: "What is the overall sentiment of the author towards pi?", criteria: { positive: "Appreciative, happy, constructive praise", neutral: "Matter-of-fact report or request without emotion", negative: "Frustrated, annoyed, upset, or angry", }, }, frustration: { type: "score", instructions: "How frustrated is the reporter?", criteria: ["not at all", "mildly", "clearly frustrated", "very angry"], }, kind: { type: "choice", instructions: "What kind of issue is this?", criteria: { bug: "Bug report or regression", feature: "Feature request or enhancement", question: "Question or support request", other: "Docs, discussion, meta, spam", }, }, }, }); if (res.stopReason !== "stop") { return { n: issue.number, title: issue.title, error: res.errorMessage }; } return { n: issue.number, title: issue.title, ...res.answers }; })); store("sentiment_results", results); return results .filter(r => !r.error) .sort((a, b) => b.frustration.score - a.frustration.score) .slice(0, 12) .map(r => `#${r.n} ${r.frustration.score.toFixed(2)} [${r.kind.choice}] ${r.title}`);Note how in that above example we also call
store()which dumps the result of that execution into the session transcript. A future invocation of Codemode can thus read back that result if it wants to.The
Promise.allhere is fine, because Pi limits the total number of concurrent tool executions itself to four and maintains a queue for the rest.A more adventurous example is to use Jev to drive a game engine for debugging purposes:
Codemode with Jev for Game Debugging
Here it knows about my
tankctlcommand and it built itself quickly a minimal harness around it to drive a game loop to assist a user with debugging a problem. Note how it built a 30 step loop in which each step goes back to both the game engine to get a text dump of what's going on, and then to Jev to determine what to do next:const jev = await models.getModelOfType("classifier", "typesafe", "jev-latest"); const tank = async (cmd) => (await tools.bash({ command: `tools/tankctl "${cmd}"` })).output; await tank("start --map assets/maps/night_arena.map"); const questions = { action: { type: "choice", instructions: "You control the tank '@' in a top-down tank game. " + "Choose the best next action.", criteria: { attack: "an enemy has line of sight to you and you can fire at it", approach: "no enemy has line of sight; drive toward the nearest enemy", dodge: "an enemy shot is heading at you and will hit soon", powerup: "a powerup is close and no enemy threatens you", }, }, }; function commandFor(choice, st) { const p = st.player; const enemy = st.enemies.filter(e => !e.dead) .sort((a, b) => (b.los - a.los) || (a.dist - b.dist))[0]; if (choice === "attack" && enemy) { return `fire_at tank ${enemy.id}; frames 30 until clear,damage,kill`; } if (choice === "dodge") { // move perpendicular to the closest incoming shot const s = st.projectiles.filter(s => !s.yours) .sort((a, b) => a.eta - b.eta)[0]; const dir = s && Math.abs(s.vel[0]) > Math.abs(s.vel[1]) ? (p.pos[1] > s.pos[1] ? "+down" : "+up") : (p.pos[0] > (s ? s.pos[0] : 0) ? "+right" : "+left"); return `input ${dir}; frames 20 until damage; input stop`; } const powerup = st.powerups.filter(u => u.available) .sort((a, b) => a.dist - b.dist)[0]; if (choice === "powerup" && powerup) { return `goto ${powerup.pos[0]} ${powerup.pos[1]} 180`; } return enemy ? `goto ${enemy.pos[0]} ${enemy.pos[1]} 90` : null; } const log = []; for (let step = 0; step < 30; step++) { const st = JSON.parse(await tank("state")); if (st.state !== "playing") break; const threats = st.projectiles .filter(s => !s.yours && s.miss_dist < 1.5 && s.eta < 1.5) .map(s => `incoming shot dist ${s.dist} eta ${s.eta}s`) .join("\n") || "no incoming shots"; const r = await models.classify(jev, { state: { map: await tank("view 8"), threats, hp: st.player.hp }, questions, }); if (r.stopReason !== "stop") { log.push(`#${step} classifier error: ${r.errorMessage}`); break; } const choice = r.answers.action.choice; const cmd = commandFor(choice, st); if (!cmd) break; log.push(`#${step} hp=${st.player.hp} ${choice} -> ${await tank(cmd)}`); } return log.join("\n");Calling MCP Servers
Lastly, Codemode obviously is great for calling MCP servers. And because we do not actually expose any of the MCP tools to the LLM, the agent first uses provided APIs to issue a tool search within Codemode to discover what it might be able to do with the connected servers. This form of progressive discovery makes the whole MCP business work well enough for a lot of use cases today.
Here for instance you can see the agent reach for the Sentry MCP straight away, even without discovering the tools, presumably because it has learned during the RL process already about what the Sentry MCP looks like. But it learns from what we inject into the system prompt, that the Sentry server is available to begin with. It's not completely guessing here.
const orgs = await tools.mcp__sentry__find_organizations({}); const { organizations } = orgs.structuredContent; const results = await Promise.allSettled(organizations.map(org => tools.mcp__sentry__find_projects({ organizationSlug: org.slug, regionUrl: org.regionUrl, }) )); return organizations.map((org, i) => { const r = results[i]; if (r.status !== "fulfilled") return { org: org.slug, error: String(r.reason) }; if (r.value.isError) return { org: org.slug, error: r.value.content }; return { org: org.slug, projects: r.value.structuredContent.projects.map(p => p.slug), }; });Modern MCP Is A Fight
I really don't want to talk too much about MCP here, but MCP is in fact a protocol that greatly benefits from Codemode. The problem in parts is that MCP in practice often targets harnesses that do not (yet?) use Codemode. But the tide is shifting. In the meantime, a temporary crutch has been to do what Cloudflare did, and do Codemode within the MCP server. But now we have Codemode in Codemode which is pretty bad. It means double JSON escaping, easy for smaller models to get confused by and the inner code cannot call the outer tools. So if you for instance use the Cloudflare MCP servers in Pi, the agent needs to write JavaScript and funnel it through more JavaScript. This is really not optimal, but it's also understandable that this is happening:
const accRes = await tools.mcp__cloudflare__execute({ code: `async () => { const r = await cloudflare.request({ method: "GET", path: "/accounts" }); return r.result.map(a => ({ id: a.id, name: a.name })); }`, }); const accounts = JSON.parse(accRes.content.map(c => c.text).join("")); const out = []; for (const account of accounts) { const r = await tools.mcp__cloudflare__execute({ account_id: account.id, code: `async () => { const r = await cloudflare.request({ method: "GET", path: \`/accounts/\${accountId}/workers/scripts\`, }); return r.result.map(s => ({ id: s.id, modified: s.modified_on })); }`, }); out.push({ account: account.name, workers: r.content.map(c => c.text).join("") }); } return out;MCP Desires
So to end things off: how well does Codemode work with MCP today? Well … not amazingly well. That's because MCP servers are not really targeting harnesses that use Codemode yet (though at this point I think most harnesses support it).
For this to work well some recommendations:
- Structured content: Codemode wants calls to return some nicely formatted JSON. So that needs to come back from the server, and many don't do that yet. The
outputSchemasystem in MCP is great for that. - Consistent results: an interesting failure case is when an MCP server does not return consistent data. For instance because it tries to token optimize things depending on how many items are in the result set. This can cause an initial probe with 5 items to succeed, but then fail when the server returns the maximum batch size.
- Large binary data: today MCP does not yet support large binary data so quite a few use cases that are really interesting do not work well at all yet. You end up with all kinds of weird workarounds such as pre-signed URLs to allow file uploads then to happen through non MCP channels.
- Composable tool search: the MCP server might know better than the MCP client which tool is appropriate for a task. But there is no good mechanism today that allows a harness to fan out tool searches across multiple MCP servers. It's all emergent behavior and it does not scale well to multiple active servers.
Future of Codemode
So where does this leave us? Is this a reversal of what I wrote a year ago where I encouraged CLIs? I don't think so. In fact, the MCP ecosystem from my perspective picked up on exactly what we pointed out a year ago works: code. But Codemode goes beyond MCP in that it can act as a capable mechanism within the harness to express more freedom for the agent.
There are however also some things that we still need to figure out. For one, durability with Codemode is trickier. We might have to adopt some ideas from durable workflow engines here to snapshot invocations. Or maybe, something like Starlark is a better composition language than JavaScript given its deterministic nature.
Images, binary data and just the inability of this pattern to work with smaller models is also something that needs to be fleshed out. So it's for sure not a perfect solution yet, but it's quite a useful pattern that I expect us to leverage more.
- Structured content: Codemode wants calls to return some nicely formatted JSON. So that needs to come back from the server, and many don't do that yet. The
-
🔗 Ampcode News Many, Many Pucks rss
You can now have multiple separate conversations with Puck.
We know you love Puck. We also know you've been asking Puck about your recent orbs, a bug fix, and dinner plans, all in the same conversation. Now each of those can be its own conversation, and each one gets its own Puck.

Press + to start a new conversation. Click the conversation's name at the top to switch between them. Conversations you haven't touched in three days move into Inactive so the list stays short.
Read more about Puck conversations in the docs.
-