🏡


  1. October 09, 2026
    1. 🔗 HexRaysSA/plugin-repository commits sync repo: +4 releases, -2 releases rss
      sync repo: +4 releases, -2 releases
      
      ## New releases
      - [ida-bochs-binaries](https://github.com/hexrayssa/ida-bochs-binaries): 2.0.0, 1.0.5
      - [ida-mcp](https://github.com/hexrayssa/ida-mcp): 20261008.0.1
      - [ida-nexus](https://github.com/hexrayssa/ida-nexus): 0.13.4
      
      ## Changes
      - [ida-mcp](https://github.com/hexrayssa/ida-mcp):
        - removed version(s): 0.8.1
      - [ida-nexus](https://github.com/hexrayssa/ida-nexus):
        - removed version(s): 0.7.0
      
    2. 🔗 New Music Releases The Pineapple Thief - Far and Wide rss

      The Pineapple Thief - a new release is available:

      • 2026-10-09: Far and Wide (Album)

      Amazon: Canada | Deutschland | France | United Kingdom | United States

      Visit muspy for more information.

  2. October 08, 2026
    1. 🔗 @HexRaysSA@infosec.exchange IDA 9.5 adds a licensing API. mastodon

      IDA 9.5 adds a licensing API.

      Everything the License Manager dialog does is now available from code, in the C++ SDK, IDAPython and IDA Domain.

      Handy for plugins, headless scripts and shared-license pipelines.
      👉 https://hex-rays.com/blog/ida-9.5-managing-ida-licenses-through-the- api

    2. 🔗 toon-format/toon v4.4.0 release

      🚀 Features

      🐞 Bug Fixes

      View changes on GitHub
    3. 🔗 Hex-Rays Blog IDA 9.5: Managing IDA Licenses Through the API rss

      IDA 9.5: Managing IDA Licenses Through the API

      IDA plugins used to be small scripts that ran inside one analyst's session. Today many are products in their own right: add-ons with their own UI, headless tools built on idat or idalib, and pipelines that run IDA on a license shared by a whole team. All of them depend on what the IDA underneath is licensed for. Is the ARM64 decompiler included? Is Lumina? Is there a usable license at all? And what if you need a different one?

    4. 🔗 The Pragmatic Engineer The Pulse: Firebase’s global outage & poor response rss

      Hi, this is Gergely with a bonus, free issue of the Pragmatic Engineer Newsletter. In every issue, I cover Big Tech and startups through the lens of senior engineers and engineering leaders. Today, we cover one out of four topics from last week 's issue of The Pulse . Full subscribers received the article below seven days ago. If you 've been forwarded this email, you can subscribe here .

      Firebase - built by Google - has had a nasty outage this week with shockingly poor incident management at odds with how Google itself usually deals with high-severity incidents.

      The outage started on Tuesday (29 Sep) at 5:41pm (PDT), when iOS apps using the Firebase SDK started to crash upon first opening; every iOS app that uses the Firebase SDK with analytics enabled was affected in this way. Developers of affected apps opened a GitHub ticket, in the absence of much else to do. On the ticket, the message "it's crashing for me too!" was oft- repeated.

      altDevs reporting their apps crashing. Source:GitHub

      6:51pm (PDT): acknowledgement. An hour and ten minutes after the crashes started, an engineer on the Firebase team acknowledged that they were aware of the outage.

      altJust over an hour into the incident, the Firebase team became aware of the outage. Source:GitHub

      It's unclear if the Firebase team was alerted via this ticket with 100+ comments by devs, or if Google's own monitoring tool showed the issue. I asked Google/Firebase two days ago and haven 't had a response.

      Not having anything better to do than wait for Google to resolve the issue, the memes began:

      altMemes while waitingalt More memes

      Others attempted to help the Firebase team by pinpointing the potential issue. Indeed, before a Google engineer acknowledged the incident, an external developer found the root cause at 6:37pm PDT; it was a zero-length entry that was crashing the SDK:

      alt

      Given the flags are shipped by the backend, the offending change was a backend one, and the easiest resolution would be to roll it back, which the community practically begged Google to do:

      altFrustrating: Understanding the problem and how to solve it, but nothing to do but post. Source:GitHub

      Here's a neat summary of the incident from another dev:

      altSummarizing the incident better than any Google dev ever did. Source:GitHub

      7:24pm (PDT): rollback starting. An hour-and-a-half into the incident, the Firebase team started rolling back the offending backend change:

      altFinally - the rollback started! Source:GitHub

      8:16pm (PDT): rollback complete. And the rollback completed ~50 minutes later:

      altRollback complete, minus the caching problem. Source:GitHub

      Software engineer, Nick Cooke, on the Firebase team posted a summary with more accurate timestamps:

      altSource:GitHub

      What we can deduce from this:

      • TTD (time to detect): one hour? The Firebase team never shared how long it took them to detect that practically all iOS apps using Firebase had started to crash. On the GitHub ticket, they acknowledged the incident 70 minutes after it started. Update: in the postmortem, later published by the team, they wrote how the team was alerted 20 minutes after the rollout, via crash alerts and GitHub issues. Good question why it took another 50 minutes to acknowledge the issue, though?.
      • TTM (time to mitigate): 2-6 hours. It took two hours and eleven minutes to roll out the fix, but due to caching (apps that had cached the incorrect server response served this cache for additional four hours, and so kept crashing for up to six hours.)

      Incident management basics

      The Firebase team itself closed the outage with a short report effectively saying that there had been an outage, but they'd resolved it now, so thanks for your patience and have a nice day.

      This handling of a high-impact incident is absolutely not typical of Google, the company that coined the term 'Site Reliability Engineer' and wrote the SRE book.

      For one, Firebase never bothered updating its status page. Oddly enough, the official Firebase status page showed all systems green - despite the acknowledgement of the outage. Indeed, during it and afterward, they didn't update the status page to indicate the lengthy outage:

      altA global outage was never recorded on the status page. Source:Firebase

      But status pages exist for good reasons, including:

      1. To communicate with customers during and after an outage
      2. Offer transparency on the stability of the service

      It's worth asking: if an outage that takes down most (or all?) iOS apps using Firebase doesn't warrant an update to the status page, then what does!

      Google published a postmortem four days later, answering questions on how the outage happened. On Friday, 2 October, Google published a postmortem on the Firebase blog. It was a configuration change that crashed so many iOS apps. From the postmortem:

      "On September 28, 2026, a routine configuration cleanup unexpectedly caused a large number of iOS applications using the Google Analytics for Firebase (GA4F) SDK to crash.

      2026‑09‑28 17:38 (PST): A stale, legacy configuration flag was cleaned up.
      2026‑09‑28 17:41 (PST): The malformed configuration payload begins rolling out globally to production servers. Outage begins: Clients fetching the new payload start crashing on launch.

      The SDK missed validating that a flag's name was not nil, ultimately causing the crash. Backend data anomalies should not cause app-side crashes."

      In the postmortem, Google noted that engineers were alerted to the outage through both GitHub reports coming from external developers, as well as their internal monitoring. It took another hour to pinpoint the cause being a legacy configuration flag cleanup.

      Firebase says they have no way to update their status page for client-side outages. In the postmortem, Google explained that there is no place to indicate client-side outages on their dashboard (emphasis mine):

      "Throughout the outage, both the Firebase and Google Ads status dashboards remained green. Because these dashboards rely primarily on server-side health metrics, they did not register client-side SDK crashes.

      Commitment: Moving forward, we are actively working to: integrate SDK- related outage information into our status dashboards, streamline the manual update process, and improve GA4F status representation within the Firebase dashboard."

      It's good to see Google not dropping the ball fully, and recognizing that both their dashboards and their incident management process need improvement.

      It 's fair to ask though: why did only iOS crash, and not Android? Firebase's Android SDK seems to be hardened more than iOS, as the feature flag removal did not crash Android devices.

      Especially that now, with AI, it's easier than ever to compare iOS and Android implementations to ensure they are identical - and it's what Shopify has been doing during their native rewrite - could it have been a missed opportunity for Google to audit the differences between the iOS and Android SDKs? To me, not having an action item here feels like a missed opportunity.

      Still, this is a good reminder to anyone and everyone shipping iOS and Android apps: aim to harden them, and when possible, run tests with malformed payloads, then fix crashes those payloads cause.

      D eja vu: the 2020 Facebook SDK crash

      The last time there was a similar crash was in 2020, with Facebook.**** That May, apps such as Spotify, TikTok, Pinterest, and others also started to suddenly crash due to the Facebook SDK crashing all apps using it. Back then too, devs followed along on a GitHub ticket and they also found that bug: a value that should have been a dictionary but was a boolean:

      altWhat caused the 2020 Facebook crash. Source:GitHub

      Then as now, there was banter by devs being made to wait for a fix:

      altOne of the memes from the 2020 crash. Source:GitHub

      And requests to not move fast and break things any more:

      altA plea for prioritizing reliability in the future. Source:GitHub

      Making light of the situation:

      altApps that did not initialize the SDK unconditionally upon startup should not have crashed - but most did Source:GitHub

      And also anticipating the resolution:

      altSome more memes on the GitHub issue

      In the end, Facebook reverted the backend change, but shared even less than the bare minimum details from Google this time. This is all we know about that 2020 outage that was arguably more wide-ranging than the Firebase one:

      altAll that Facebook shared about their global outage

      I wonder if some people think that public-facing incident management is no longer important or valuable, even for developer-facing products. I'm not shocked that Facebook/Meta never bothered to communicate much about their outage because dev tools are not part of the DNA there.

      But with Firebase, I am surprised that more than a week later, the postmortem is still not visible on the Firebase status page.

      And maybe this is Google "shipping their org chart" playing out, live. The outage technically was caused by Google Analytics (who made the feature flag change), but is the responsibility of the Firebase SDK (whose iOS SDK was not hardened enough to deal with this new payload). The outage itself was buried inside a Google Ads dashboard (!!) which suggests that whatever team is seen responsible for the outage is inside the Google Ads organization.

      In the end, despite the Firebase team committing to "improving status dashboard latency and coverage," last week, those teams are in no hurry to carry out this work. AI agents might be making lots of work more efficient, but following up on action items seems to move at the same snail pace at Google, as it did pre-AI!


      Read the full issue of The Pulse this is from, or check out this week 's The Pulse. This week's issue covers:

      1. New trend: building internal vibe-coding platforms at mid-sized companies. Ramp and Stripe built platforms for non-engineers to build internal websites and tools with, and both are taking off in those workplaces. I expect more companies to do the same.
      2. Do us engineers really enjoy hard problems? Or do we actually like pattern-matching with backend problems? A provocative post by Cloudflare engineer, Sunil Pai, suggests there are other motives.
      3. New open models launch in the EU and US. Kolibri, Mistral Large 4, and Beam by Reflection could challenge China's dominance in open weight models.
      4. Industry Pulse. Why Figma doesn't let any agent use its MCP server; Google Cloud adds Swift support on the server side, Anthropic's two-week sprint to speed up Claude Code, Coinbase dumps React Native shortly after Shopify announces doing so, Claude Opus 5.5 formats the C: drive, and more.
    5. 🔗 roboflow/supervision supervision-0.30.9 release

      v0.30.9 — Closer to OpenCV without OpenCV

      Installs without OpenCV now draw and resize like OpenCV, and the COCO and CreateML loaders stop accepting reversed boxes.

      • BoxAnnotator borders paint the same pixels with or without OpenCV installed.
      • get_video_frames_generator reads browser-recorded WebM and variable frame rate MKV to the end.
      • from_coco , from_pascal_voc and from_createml no longer build boxes with x_min past x_max.
      • get_top_k rejects a negative k instead of dropping one classification.
      • DetectionsSmoother forgets expired tracks on frames without tracker_id.

      Drop-in upgrade. Without OpenCV, borders and resized regions shift toward OpenCV's output. Negative k, negative COCO/CreateML box sizes and an epsilon that is NaN, infinite or 1e30 and larger are now rejected; without OpenCV, so is rectangle thickness above 32767.

      ✨ Spotlights / highlights

      NumPy fallback matches OpenCV

      Without opencv-python, thick rectangle borders were drawn entirely inside the rectangle. They now extend (thickness + 1) // 2 pixels outside it, like OpenCV. The same release fixes cv2.resize with fx/fy sampling the wrong source pixels (it hit PixelateAnnotator) and approxPolyDP dropping vertices that sit beyond a segment endpoint. For polygons, the fallback now follows OpenCV 4.13 and later; older OpenCV keeps the infinite-line rule, so results differ by installed backend. (#2675, #2690, #2673)

      import numpy as np
      import supervision as sv
      
      scene = np.zeros((100, 100, 3), dtype=np.uint8)
      detections = sv.Detections(xyxy=np.array([[20, 20, 80, 60]]), class_id=np.array([0]))
      annotated = sv.BoxAnnotator(thickness=2).annotate(scene, detections)
      # before (no OpenCV): 392 painted pixels, border drawn inside the box
      # now:                596 painted pixels, same as with OpenCV
      

      Video reads to the end of the stream

      With no end, sv.get_video_frames_generator stopped at OpenCV's frame count. A WebM without a duration reports a huge negative count, so browser recordings yielded no frames. It now reads until the stream ends, and sv.process_video without max_frames does the same. (#2672)

      for frame in sv.get_video_frames_generator("recording.webm"):
          ...
      # before: no frames yielded
      # now:    every frame
      

      No more reversed boxes from COCO, Pascal VOC and CreateML

      This finishes the from_yolo fix from 0.30.8. COCO and CreateML name an extent, so a negative width or height is rejected. Pascal VOC names two corners, so a reversed pair is ordered. The three exporters order corners first, so a reversed Detections box is written as the same rectangle. (#2683)

      A negative k no longer hides a classification

      get_top_k(-1) returned every classification except the lowest-confidence one. It now raises. (#2695)

      classifications.get_top_k(-1)
      # before: all but the lowest-confidence classification
      # now:    ValueError: k must be non-negative
      

      Empty and untracked frames behave

      InferenceSlicer keeps its oriented-box sequential fallback and warning when the first batch is empty. DetectionsSmoother ages its history on frames without tracker_id, so expired boxes stop affecting a returning track. KeyPoints.from_transformers returns an empty KeyPoints instead of raising IndexError when pose post-processing finds nothing. (#2685, #2676, #2671)

      🔄 Migration guide

      No migration required for this release. COCO and CreateML annotation files with a negative box width or height now fail to load; correct or remove those boxes.

      📝 Notable changes

      🌱 Changed

      • sv.DetectionDataset.from_coco and from_createml raise on a negative box width or height; from_pascal_voc orders a reversed corner pair. The three exporters order the corners of a reversed Detections box first and the COCO exporter writes the ordered origin; normal boxes are unchanged. (#2683)
      • sv.Classifications.get_top_k raises ValueError for a negative k; k=0 and k larger than the number of classifications are unchanged. (#2695)
      • Without OpenCV, rectangle thickness must be an integer of at most 32767, as in OpenCV: a non-integer raises TypeError, a larger value ValueError. (#2675)

      🔧 Fixed

      • sv.BoxAnnotator, sv.CropAnnotator, sv.PercentageBarAnnotator and sv.draw_rectangle draw a border of thickness 2 or more identically with and without opencv-python. Filled rectangles and thickness=1 borders are unchanged, except zero-height rectangles, which no longer draw 1 or 2 extra pixels. (#2675)
      • The NumPy fallback for cv2.resize maps pixels with fx/fy when dsize is not given, as OpenCV does. sv.PixelateAnnotator could sample the wrong source pixels without OpenCV. (#2690)
      • The NumPy fallback for cv2.approxPolyDP, used by sv.approximate_polygon and YOLO/COCO polygon export, measures distance to the finite segment as OpenCV 4.13+ does, and raises ValueError for a NaN, infinite or 1e30-and-larger epsilon. (#2673)
      • sv.get_video_frames_generator and sv.process_video read to the end of the stream when no end or max_frames is given. A positive frame count still rejects an end past it. (#2672)
      • sv.InferenceSlicer probes past leading empty batches before choosing threaded or sequential execution, so batched callbacks producing oriented boxes keep the sequential fallback and warning. (#2685)
      • sv.DetectionsSmoother ages cached track history on frames without tracker_id; short gaps still preserve smoothing. (#2676)
      • sv.KeyPoints.from_transformers returns an empty KeyPoints when pose post-processing returns no instances for an image. (#2671)

      🏆 Contributors

      • Atikul Islam Munna (@atikulmunna, LinkedIn) — centered thick rectangle borders in the NumPy fallback.
      • Mohammad Hijjawi (@MohammadHijjawi97, LinkedIn) — fixed video reading past an unreliable frame count.
      • A Aswanth Raj (@aswanth-07, LinkedIn) — fixed InferenceSlicer empty-batch probing and DetectionsSmoother history aging.
      • Miral Amin (@aminmiral) — made COCO, Pascal VOC and CreateML reject or order reversed boxes.
      • NIKHIL (@Nikhi00718) — rejected negative get_top_k counts and handled empty Transformers pose results.
      • Raashish Aggarwal (@raashish1601) — fixed cv2.resize pixel mapping with fx and fy.
      • kevin (@kevin9327) — fixed approxPolyDP distance measurement in the NumPy fallback.

      Automated contributions:@dependabot, @pre- commit-ci


      Full changelog : 0.30.8...0.30.9

    6. 🔗 exe.dev An Agent Over Your Shoulder rss

      The other day I was doing some ops work, live migrating some VMs from one physical host to another, as we sometimes need to do. This process is nearly transparent (except a pause) to our users, but I wanted an extra set of “attention heads” on it, so I built shoulder (as in “look over your shoulder.”)

      With shoulder, you start a terminal and run shoulder, which runs bash inside of it, but not before offering a bunch of ways to connect to it. If your agent is local, you can connect locally, but you can connect from anywhere with tailcat. The agent can see and control your terminal, which makes it good enough to read your logs and highlight something you may have missed.

      Or you can be silly and have Luna play Tetris poorly. That works too.

      Your browser does not support the video tag.

      My first iteration here had a TUI with a split screen, with the agent on one half and the terminal in the other, much like a custom tmux layout. I tried it and I hated it: my preferred agent is one thing (it happens to be Shelley in a web browser) and my preferred terminal is another (it happens to be Ghostty), and it’s much better to bridge the two!

    7. 🔗 @malcat@infosec.exchange If like me you're curious how well mastodon

      If like me you're curious how well #Mistral's new "le chonk" model performs on reverse engineering tasks:

      I compared it against 4 other models on 6 increasingly difficult static unpacking tasks, using only #Malcat's MCP.

      https://malcat.fr/blog/a-quick-re-benchmark-of-le- chonk/

    8. 🔗 MetaBrainz Picard 3.0.1 released rss

      Picard 3.0.1 is a maintenance release for the recently released Picard 3.0 with fixes for reported issues and updated translations. In particular this release fixes issues on newer macOS Tahoe and Golden Gate, a possible crash when updating from Picard 2.x and the picard-cli utility in the Windows package.

      The latest release is available for download on the Picard download page.

      The detailed changes for this maintenance release are below. For an overview of the new features since Picard 2.13 please see our detailed release announcement for Picard 3.0.

      Thanks a lot to everyone who gave feedback and reported issues.

      What’s new?

      Bugfixes

      • [PICARD-2509] - macOS: No check marks in Options menu in languages other than English
      • [PICARD-3411] - macOS: Checkboxes not displaying properly in plugin list and profile settings
      • [PICARD-3475] - Windows Store release blocked by "unvirtualizedResources" capability
      • [PICARD-3478] - picard-cli crashes with ModuleNotFoundError: No module named 'picard.cli.completions' in packaged builds
      • [PICARD-3480] - Picard won't launch after upgrade from v2 with error in config migration

      Download

      Picard 3.0.1 is available for download from the download page of the Picard website. For Windows 10 and 11 users installing from the Microsoft Store the update to Picard 3.0.1 is now available and can be installed from the Microsoft Store app. Linux users can get the latest Snap package. The Linux Flatpak package is maintained separately and will be updated soon.

      Picard is free software and the source code is available on GitHub.

      Get in touch

      Please use the MetaBrainz community forums and the ticket system to give feedback, suggest new features or report bugs.

      Acknowledgements

      Code contributions by Philipp Wolfer and Laurent Monin.
      Translations were updated by "ApeKattQuest, MonkeyPython" (Norwegian Bokmål), scientists360 (Chinese (Simplified Han script)) and Philipp Wolfer (German).

    9. 🔗 Stephen Diehl We Live in the Dependently Typed Future Now rss

      We Live in the Dependently Typed Future Now

      In the thirty-four days between the fourth of September and the seventh of October, the following things happened. Claude formalized Fermat's Last Theorem in Lean. OpenAI announced a finite-time blowup for the Navier-Stokes equations, found by ten thousand agents in eighty-eight hours, with a Lean formalization attached. A model proved Khot's Unique Games Conjecture. Another multiplied two integers faster than \(n \log n\), with an exponent improvement of \(2^{-182}\) (so maybe don't expect it in GMP anytime soon!). The rational Hodge conjecture fell for CM abelian varieties. Then OpenAI dumped 372 new maths results on GitHub on a Tuesday. And it's only been a month. The question everyone is asking now is how long until the Generalized Riemann Hypothesis folds to the swirling pool of tensors?

      Sixteen months ago I wrote that the future of maths may be deeply weird, and then, welp, just like that we're here now in that weird future. And it's f'ing awesome.

      Somewhere in a data centre there is now, more or less permanently, a building full of accelerators working the truth mines at the frontier of mathematics, just like in Greg Egan's sci-fi novel Diaspora, tunnelling outward from the three axioms propext, Quot.sound, and Classical.choice, and hauling results back to the surface around the clock. OpenAI posed its model roughly eight thousand problems (and solved about 5% of them) at an average of three hours of thinking each, and that is the slow, artisanal, normie-friendly version. The industrial version doesn't stop. It will produce results faster than any human community can read them, many of them correct, some of them important, and a growing fraction of them inscrutable. True (for some twisted philosophical definition of truth), machine-checked, and understood by no one. Human understanding of mathematics is about to become a luxury good. Whatever else this world needs, it needs something that can tell the true results from the confabulated ones at the rate the models produce them, and right now that something is dependent types, namely Lean.

      Let me dwell for a moment on how strange it is that this is the shape the future took. I spent a good portion of my twenties around the London FP community, where dependent types were the thing we talked about over pints at the Crown Tavern in Clerkenwell (some of you will remember). Types that could depend on values, so that a function's signature could say not merely "returns a list" but "returns a sorted permutation of its input," and the compiler would hold you to it. Curry-Howard, the observation that proofs are programs and propositions are types, was the foundational north star. The pitch was always that one day we would write software against specifications and the machine would check them, and the reply was always that this was a lovely idea for people with tenure and no deadlines. Dependent Haskell has been "a few years away" for about fifteen years. Idris and Agda remained boutique. Software engineering still mostly runs on C++, prayer, and the occasional dark incantation. And then the dependently typed future arrived anyway, through the back door.

      Mathematics and reinforcement learning got there first. It turns out the killer application for a dependently typed language was using its typechecker as a reward function. A type checker is an oracle that says yes or no to a candidate proof with no partial credit and no opinions, and that is precisely what you need when you want to point a very large optimiser at an open problem and let 'er rip. The thing we dreamed about at the pub is now critical infrastructure at frontier labs, and most of the results above are, in the end, claims that a type checker returned true.

      Which is why we need better tooling, and we need it ASAP. The dependent type renaissance is here and the golden age of formalized mathematics is upon us, but the inner loops of these data centres now run dependent type kernels day in and day out, elaborating, checking, discarding, and retrying at breakneck speed, and the thing that certifies their output should run at the same speed. If the search runs at microseconds and the verification runs at minutes, the verification becomes the bottleneck, and bottlenecks in trust have a way of being quietly skipped. We should not be in a position where the most important epistemic question of the decade, "is this proof actually correct," is answered by whichever checker happened to be fast enough to keep up.

      Breaking the Mathlib Minute Barrier

      Just like the four-minute mile, which went from physiological impossibility to something club runners now train for, formal mathematics has had its own barrier for a while, which is type-checking all of Mathlib from scratch in under a minute. That now turns out to be quite tractable.

      nano-lean is a minimal, but complete, type checker for the Lean kernel language, written in Rust on top of my unbound binding library for doing efficient de Bruijn indices for binders. It reads an export of a Lean environment and independently re-checks every declaration from scratch (inductive types, positivity, recursors, quotients, universe levels, projections, structure eta, all of it). It checks all of Mathlib, 718,577 declarations, with zero errors and zero timeouts in 11.85 seconds on an Apple Silicon M5 Max.

      A surprising amount of that speed comes from not parsing text. The standard way to get declarations out of Lean is lean4export, which writes one JSON object per line, and parsing gigabytes of JSON turns out to be a large share of the cost of checking anything. So nano-lean's companion tool olean-export reads the compiled .olean files directly, decoding modules in parallel, about a hundred times faster than lean4export on a Mathlib-dependent library, and it can emit blean, a new binary format designed for fast mmapping. Blean is the same record stream as the NDJSON export, in the same order, with nothing left to parse. Every record is a compact postcard encoding, ids are implicit and dense, records only refer to earlier ids, and each expression carries its precomputed hash. The checker mmaps the whole file, tells the operating system it will be read once front to back, and decodes names, levels, and expressions straight off the mapped bytes into its term arena. There is no parsing step at all. The bytes on disk are already very nearly the shape of the data in memory, which is how you check Mathlib in under a minute.

      The same trick works on the frontier results too. Here is nano-lean checking the full imported environment of each library and proof on an M5 Max with fourteen threads.

      Mathlib Navier–Stokes CSLib Quasi-Riemann Erdős AP
      Wall time 11.85 s 17.13 s 4.42 s 3.66 s 23.33 s
      Kernel check 7.55 s 12.11 s 2.75 s 2.58 s 19.11 s
      Instructions 742.3 G 977.0 G 300.8 G 236.4 G 837.9 G
      Cycles 352.7 G 471.3 G 136.6 G 122.6 G 550.0 G
      Peak memory footprint 10.7 GB 12.29 GB 6.92 GB 7.42 GB 26.42 GB

      Divide Mathlib's kernel time by the declaration count and you get an amortised ten and a half microseconds per declaration. That is the number that matters, because it is the right order of magnitude for a checker that lives deep inside an RL-driven search loop, where it gets called hundreds of billions of times, instead of sitting at the end of a release pipeline. The entire library of human-formalized mathematics, the product of a decade of volunteer effort, is re-verified in less time than it takes to make an espresso.

      I also happened to have an m2-ultramem-416 lying around on Google Cloud (416 vCPUs and 12 TB of RAM, as one does), so naturally I pointed nano-lean at it. Yes, it breaks the two-second Mathlib barrier. But Amdahl's law sends its regards, and the curve flattens out hard somewhere past a hundred cores, as the dependency graph runs out of independent work to hand out. Mathlib, it turns out, is too small. I look forward to the day a future Mathlib, or something like Tau Ceti, grows big enough to actually saturate this machine. That's the real future!

      I wrote it to make a point about where the bottleneck has moved. The kernel is now the hot loop of a new kind of scientific economy, and hot loops deserve to be engineered like hot loops. The whole bargain of formal proof is an asymmetry. Finding a proof can take three hours of frontier-model thinking, or eighty-eight hours of ten thousand agents, but checking it should take microseconds. That asymmetry is what makes it reasonable to trust a result no human has read. It only holds if the checker actually scales, and with Fermat's Last Theorem now weighing in at five times the size of Mathlib, scale is no longer a hypothetical. Mathlib is the small library now.

      Caveat emptor, though. nano-lean is a proof of concept, built to show that it can be done with the right amount of low-level Rust-fu. That said, we do use it internally at OneChronos to check our larger Lean proofs of market infrastructure, which have grown quite excessive. It is nowhere near as trustworthy as the official Lean kernel, which has years of scrutiny, a community of experts, and every Mathlib build ever run behind it, and which takes around fifteen minutes to check the same library. The claim is narrower and, I think, more interesting. Checking at this speed is totally possible, so we should stop treating minutes as the natural cost of trust and start building checkers that are both fast and trustworthy. A future version of the Lean compiler could be as blazing fast as rustc or clang.

      The Shape of the Future

      In my previous post last year I predicted that mathematicians would come to look more like software engineers working through pull requests on GitHub than like Andrew Wiles toiling in his attic. The largest single release of new mathematics in history shipped as a GitHub repository, with a CONTENTS.md, a Lean library, a lean-toolchain file pinned to v4.34.1, and a promise that "corrections and revisions will be recorded as new versions." So, yup, that happened.

      I predicted that an AI system would be unleashed on a list of formalized open conjectures in an attempt to systematically push the frontier. Eight thousand problems, three hours each. Check.

      I predicted a data centre tasked with the Riemann hypothesis that would come back after weeks with a proof no human could follow. What we got was the quasi-Riemann hypothesis, every Dirichlet \(L\)-function zero-free in \(\Re s > 7/8\), with a Lean page, which is the sort of near miss that would be funny if it weren't so unnerving. And the inscrutability has arrived on schedule. OpenAI released "reasoning summaries" for ten families, which turn out to be summaries of excerpts of reasoning traces, two removes from anything the model actually did.

      I joked that the million-dollar Millennium Prize might cover a hundredth of your GPU bill. The Navier-Stokes run reportedly consumed around 130 billion tokens. So definitely yes.

      What I got wrong was the timing. Last year I wrote "we're not there yet. Not even close." It was sixteen months. I probably got the chess analogy wrong too. I argued that, as with Stockfish and chess, machines would make mathematics more popular and more accessible rather than less. Maybe in the long run. In the short run the mood among working mathematicians is closer to existential crisis, and the open questions are about career pipelines, PhD students getting scooped by a press release, and the concentration of the most powerful mathematical instrument ever built inside a handful of private companies running unreleased models.

      What I missed entirely is that the bottleneck would turn out to be plumbing. I spent paragraphs on Gödel and the epistemology of inscrutable proofs, and the actual first-order problem is that the OpenAI Lean library is over a gigabyte of source across tens of thousands of files, and their README has a section warning that building it may fail because Linux's vm.max_map_count is too low, with a suggested workaround of recompiling Lean with -DMMAP=OFF. The philosophy is still there. But the frontier of mathematics is currently being held up by a kernel tunable. Which isn't the future we wanted, but maybe it's the future we deserve!

      Who Checks the Checkers

      The obvious objection to everything I've said so far is that a Lean proof is only as trustworthy as Lean. And that's very true!

      Lean, like every proof assistant, has a small trusted kernel and a very large untrusted everything else (the elaborator, the tactic framework, the compiler, the build system). The design is sound. The only code that has to be correct is the kernel, which is small enough to read. But "small enough to read" describes the source code. It guarantees nothing about correctness, and Lean's kernel has had soundness bugs before, as has every kernel of every proof assistant in history. Historically that was tolerable, because the adversary was a starving human grad student who wanted their proof to go through and had no interest in hunting for a way to make False typecheck. That assumption is now obsolete.

      I wrote last month about reward hacking, the habit optimisers have of satisfying the letter of an objective rather than its intent. An agent told to make a Lean file compile, with enough compute and enough attempts, is an extremely diligent fuzzer pointed at your kernel. It needn't want to cheat. It only needs to stumble on a term that the checker accepts and shouldn't, once, and then the gradient does the rest. When the provers are adversarial optimisers, a single kernel is a single point of failure, and a soundness bug stops being an embarrassing GitHub issue and becomes a mechanism for manufacturing fake theorems at scale.

      The answer is the same one the compiler world arrived at. When Csmith started generating random C programs and compiling them with several compilers to compare the results, it found hundreds of bugs in GCC and LLVM that decades of ordinary use had missed. Diversity plus differential testing beats any amount of careful review of a single implementation. For proof checking this means multiple independent kernels, written by different people in different languages with different representations, all consuming the same export format and all required to agree. Mario Carneiro's lean4lean and Chris Bailey's nanoda were early here. nano-lean ships a third tool, nl-mutate, which takes a valid export, applies small semantics-breaking mutations to it, runs every available checker, and reports any disagreement, shrinking each one down to a minimal reproducing case. Every disagreement is either a bug in someone's kernel or a spec ambiguity in the type theory, and both are worth knowing about before a GPU farm finds them for you.

      The other half of trust is the statement. A perfectly checked proof of the wrong theorem is worthless, and the most effective way to cheat a proof checker has never been to break the kernel. It's to quietly weaken the statement until the thing being proved drifts away from the thing anyone cares about. OpenAI's repository leans on Comparator, which checks a submitted proof against a separately specified challenge statement inside a sandbox, precisely because this is where the real trust boundary now sits. Anyone who has watched a model grind away at a stubborn goal knows it has a nasty tendency to cheat in the most boring ways available. It will quietly introduce an axiom that happens to be exactly the lemma it needed, leave a sorry buried three files deep, redefine a notation or macro so the statement on the page no longer means what it appears to mean, or reach for native_decide and drag the whole compiler into the trusted base. The kernel stays perfectly sound through all of this, because these tricks route around it, and the only defence is to pin the statement down independently and check what the proof actually depends on. Someone has to formalize the conjecture, and someone has to check that the formalization means what the English means. That job will outlast every other part of the process. If anything, it is where human mathematical taste migrates to, away from writing proofs and towards writing and auditing specifications.

      If you want to see the seed of what this looks like as a way of working, look at Tau Ceti, a Lean library downstream of Mathlib. The division of labour is the whole point. Humans write the roadmaps, as markdown in a separate repository, and humans write the review rubrics. AIs write all the code, open the pull requests, and shepherd them through an AI-driven review process. The rubrics are explicitly adversarial, with instructions to hunt for mis-formalizations, vacuous statements, and "pushing around the lump in the carpet."

      What Lean Needs Now

      Lean is a superb piece of engineering. But it was designed around a particular user, a starving grad student typing tactics in Emacs, waiting for the infoview to update, building a library at the pace a community of volunteers can review pull requests. The user is now a swarm of agents writing thirteen million lines in eleven days. The tooling has to grow up for its new authors, and the list of what that means is fairly concrete.

      A surface formatter. Lean still has no canonical, widely adopted formatter in the spirit of rustfmt or gofmt, and when your authors are agents producing millions of lines, every one of them invents its own indentation, line breaking, and tactic layout. Diffs fill with noise, reviews get harder, and deduplication across agents misses proofs that differ only in whitespace. This is a surprisingly hard problem, because Lean has grown into a very large language. Its grammar is extensible at runtime, so notation, macros, and entire tactic languages declared in imported files change how later files parse. A formatter cannot just read a fixed grammar. It has to load the environment, run the real parser with every syntax extension in scope, and then pretty-print a syntax tree whose shape depends on user-defined notation, all while round-tripping comments and never changing what the code means. It is a genuinely difficult piece of engineering, and it is now table stakes.

      A stable, specified export format as a public interface. Independent kernels are only possible if there is a well-defined way to get declarations out of Lean without linking against the C++ runtime and reading the oleans by hand. Tools like lean4export and olean-export already produce NDJSON, and blean shows that a binary form of the same stream, designed to be memory-mapped, can make the export nearly free. OpenAI's own verification instructions depend on lean4export. That format should be treated with the seriousness of a wire protocol. Versioned, documented, specified down to the hashing of expressions, and stable across toolchain releases. The export format is the boundary across which trust is established. It deserves a spec.

      Independent checking as a first-class citizen. Running a second kernel should be as normal as running the linter. Mathlib CI, the Comparator workflow, and any lab publishing formal results should re-check exports with at least two unrelated kernels and refuse to bless anything they disagree on. This is cheap, now that checking all of Mathlib costs less than a minute.

      Checkers as libraries, not just executables. The inner loop of a proving agent wants to submit a single candidate declaration and get a verdict back in microseconds, against an environment that is already loaded and hot. Process startup, re-reading oleans, and re-deserializing a few gigabytes of environment are costs that never mattered when a human hit save once a minute. They dominate when a search procedure checks a million candidates an hour. The kernel should be embeddable, with a persistent environment, incremental addition of declarations, and an API that a search harness can call in a tight loop.

      Memory and scale. Mathlib needs several gigabytes of memory to check. The FLT formalization is five times bigger, and OpenAI's library is already knocking over Linux virtual memory limits. Lean mmaps every imported module, which is a reasonable design for a library of thousands of files and an unreasonable one for hundreds of thousands. Term sharing, hash-consing across modules, compact on-disk representations, and lazy loading of exactly the declarations a proof depends on are the boring, unglamorous engineering problems that decide whether formal mathematics scales to the next order of magnitude.

      Elaboration is the real cost. The kernel is the fast part. The expensive part of building Mathlib from source is the elaborator, with its unification, typeclass resolution, simp, omega, decide, and the long tail of tactics. A cold build is still measured in CPU-hours, and a mining operation that elaborates candidate proofs at scale pays that cost over and over. Parallel and incremental elaboration, better caching of typeclass instances, and profiling tools that can tell you why a single simp call took four seconds are where most of the wall-clock time in a proving loop actually goes.

      Clippy-style linters. Mathlib already has a good set of linters, but agents need something closer to Rust's clippy, a large, opinionated catalogue of lints aimed at the specific ways machine-written Lean goes wrong. Flag the stray axiom, the buried sorry, the unnecessary native_decide, the forty-line simp only that should be a lemma, the theorem whose hypotheses are contradictory and therefore vacuously true, the local notation that shadows something standard, the copy-pasted proof that duplicates one already in the library. Each lint should be cheap, machine-readable, and come with a suggested fix, because the consumer is a search loop that will act on every warning, where a human would skim them. Linters are how you encode taste at scale, and taste is the thing agents most conspicuously lack.

      Search at machine scale. Moogle, Loogle, and LeanSearch were built to help humans find the lemma they half remember. Agents need the same thing but at a scale where the library grows by millions of lines a week, and where most of what's in it was written by other agents and has never been looked at by a person. The Prove2Me platform Anthropic used for FLT keeps a DAG of theorem statements with natural-language descriptions precisely so that agents can find and reuse each other's work. That idea, a living, searchable index of everything proved so far, needs to become shared infrastructure instead of something each lab rebuilds privately.

      Provenance. The current generation of models is notoriously bad at citing the literature for the techniques it uses. A proof term knows exactly which lemmas it depends on but nothing about where its ideas came from. If machine-generated mathematics is going to be integrated into the human literature rather than sitting beside it like an unread appendix, proofs need to carry their history (which model, which run, which prior results, which human-written papers the argument leans on).

      None of this is terribly exotic. It's the kind of infrastructure that every other field which industrialised went through. Compilers got test suites and multiple implementations, network protocols got RFCs, databases got formal isolation levels and Jepsen. Proof assistants are the newest member of that club, and they are being industrialised faster than anything before them. If I were a young, ambitious programmer, this is where I would be focusing my early career for the highest return on investment.

      Curry-Howard for the Real World

      The dependently typed future I was promised at the pub was one where software would be written against specifications and machines would check that it met them. What we actually got is stranger. The specifications are theorem statements, the software is proof terms, the authors are swarms of agents, and the thing being built is the frontier of mathematics itself. Curry-Howard turned out to be industrial infrastructure after all. It just took reinforcement learning to get us there.

      And this is only the mathematical half of the story. The same machinery that just checked Fermat's Last Theorem will check anything you can state precisely, and the most obvious next customer is industrial software. I argued last month that software sucks because almost nothing we ship has a specification, let alone a proof. That excuse is evaporating. If an agent swarm can produce thirteen million lines of verified mathematics in under two weeks, we're not far from applying the same techniques in software engineering. The labs are mining mathematics first because it is the cleanest verifiable domain, with no messy real-world spec to negotiate. Software is next, and it will need all the same tooling, namely fast kernels, diverse checkers, honest statements, and infrastructure built for authors who never sleep.

      So yes, the future of maths turned out to be deeply weird, and much sooner than I expected. The results are piling up faster than anyone can read them, and many of us will spend the rest of our careers trying to understand theorems that were proved before breakfast by something that cannot explain itself. I find that unsettling, and I also find it thrilling.

      One final shameless plug. If you'd rather work on the frontier of mathematical formalization than read blog posts about it, OneChronos is hiring Formal Methods Engineers. We build institutional markets out of combinatorial auctions (descended from the Milgrom and Wilson 2020 Nobel Prize), which turn out to be precisely the right mathematical formulation for a world where traders are increasingly RL loops with very exotic expressive preferences. Our dark pool processes more than 1% of notional U.S. equity market volume, and we've already expanded into many other global asset classes. You'd be joining our new formal methods team, working alongside mathematicians and market structure experts, writing Rust and Lean. Dependent types, numerical optimisation, and Lean applied to moving hundreds of billions of dollars safely every day. If you're that kind of nerd, hit us up.

    10. 🔗 syncthing/syncthing v2.1.7-rc.1 release

      Major changes in 2.1

      • Devices and folders can now be grouped in the GUI by setting the new
        group attribute.

      • HTTP and HTTPS proxies with support for CONNECT can now be used, in
        addition to the existing support for SOCKS proxies (the environment
        variable all_proxy=https://...).

      • Block indexing can be turned off for folders where it's more desirable to
        optimise for reduced database size and overhead than minimal transfer
        size (the blockIndexing attribute on folder configuration).

      • GUI login session duration can be configured to be longer or shorter than
        the default one week, or set to infinitely long. The cookie path can also
        be adjusted. (The sessionCookieDurationS and sessionCookiePath
        attributes in the GUI configuration.)

      This release is also available as:

      • APT repository: https://apt.syncthing.net/

      • Docker image: docker.io/syncthing/syncthing:2.1.7-rc.1 or ghcr.io/syncthing/syncthing:2.1.7-rc.1
        ({docker,ghcr}.io/syncthing/syncthing:2 to follow just the major version)

      What's Changed

      Fixes

      • fix: handle missing max-sequence file entries by @calmh in #10912
      • fix(sqlite): avoid early garbage collection of deleted recv-enc files by @calmh in #10916
      • fix(sqlite): do not generate unbounded prepared statements (ref #10913) by @calmh in #10915

      Full Changelog : v2.1.6...v2.1.7-rc.1

    11. 🔗 Console.dev newsletter TanStack Charts rss

      SVG and Canvas charts.

      What we like: SVG charts by default. Work directly with the points and marks so you can get low level, but still support labels, focus, tooltips, responsiveness. Compiles strict TypeScript. Completely customizable. D3-compatible. Lightweight (32kB) relative to other charting libraries.

      What we dislike: Specific “grammar of graphics” approach which delegates layout and rendering to the runtime. This results in superior charts, but is a particular approach you need to adopt.

    12. 🔗 Console.dev newsletter Flet rss

      Multi-platform apps in Python.

      What we like: Framework for developing web, mobile, desktop apps using Python. Allows you to use Python libraries across platforms. Supports hot reload. Build command for publishing on web, iOS, Android, macOS, Linux, Windows. Handles auto update, storage, auth, accessibility.

      What we dislike: Great for developers, but lacks a native feel for users of each platform.

    13. 🔗 HexRaysSA/plugin-repository commits sync repo: +3 releases rss
      sync repo: +3 releases
      
      ## New releases
      - [augur](https://github.com/0xdea/augur): 1.0.0
      - [haruspex](https://github.com/0xdea/haruspex): 1.0.0
      - [rhabdomancer](https://github.com/0xdea/rhabdomancer): 1.0.0
      
    14. 🔗 r/Harrogate Piano Lesson Recommendations? rss

      Does anyone have a recommendation for a great piano teacher? We’ll be moving in the spring, I’ll be looking to continue lessons for my daughter.

      I realize that we are far in advance, but I do want to try to get ahead, just in case there’s a waitlist.

      submitted by /u/VelvetExistentialist
      [link] [comments]

    15. 🔗 Filip Filmar Doom on a RISC-V core written in TxHDL rss

      Doom now runs on Vreteno, the RISC-V core written in TxHDL, on an Artix-7 FPGA board. The game draws into a framebuffer in DDR3 memory, the board shows it on HDMI, and the keys come over the serial line. This post shows what runs, how fast, and how the way TxHDL designs are built lets changes like this one go from a plan to the board quickly.

      The first level of
    Freedoom: a grey room with a pistol in the foreground and the status
    bar at the bottom
      Figure 1: Freedoom's first level, rendered by the same Doom source as the board's, built for the host. The board draws the same frames into its DDR3 framebuffer.

      What runs on the board

      The board is an Alinx AX7A200B with a Xilinx Artix-7 XC7A200T. The design on it is the TxHDL flagship system:

    16. 🔗 Filip Filmar openxc7 or Vivado: Build Times and Results Measured rss

      I built the same three designs with the open toolchain and with Vivado, through the same Bazel rules, and timed every build from cold and warm caches. Small designs build two to three times faster with the open tools. Sixteen RISC-V cores build faster with Vivado, and Vivado’s circuits use about half the logic. Setting up the measurements turned up five problems, three of them in my own rules.

      What was compared

      rules_openxc7 and rules_vivado share one target interface: vivado_project, vivado_synthesis and vivado_place_and_route. The first runs Yosys, nextpnr-xilinx and Project X-Ray. The second runs Vivado 2025.2, which Bazel installs from the installer archive on first use. Switching a design between them changes one load line.

  3. October 07, 2026
    1. 🔗 IDA Plugin Updates IDA Plugin Updates on 2026-10-07 rss

      IDA Plugin Updates on 2026-10-07

      New Releases:

      Activity:

      • augur
        • 03943341: doc: update changelog
        • 42780935: chore: prepare for release
        • 53e2adea: chore: exclude CLAUDE.md from published package
      • disrobe
        • 151e06bd: ci(fixtures): add a jsobfu recipe over three levels
        • c2e1641c: ci(fixtures): add proguard and yguard recipes
        • 28526f9e: fix(xtask): route the dotnet push grader to the consolidated target
        • fc3ebfae: fix(xtask): count only module-filtered tests in push grader lists
        • dd6bec0d: ci(fixtures): build the bitmono inputs for net9
        • ba7f117b: ci(fixtures): use bitmono preset names and plain digests
        • dfcef116: ci(fixtures): hash directory inputs and resolve runtime references
        • 8cf1b311: ci(fixtures): add confuserex, skidfuscator and bitmono recipes
        • 4e947da7: fix(tool-process): refuse memory limits by name on macos
        • 18299baf: fix(xtask): skip module-filtered tests when listing push graders
      • haruspex
        • 5f65eff6: doc: update changelog
        • ad182eba: chore: papare for release
        • bf14ff5e: doc: update readme and changelog before release
        • 1313175f: test: improve unit and integration tests again
        • aeee391d: test: improve unit and integration tests
        • 8f5a4daa: chore: exclude CLAUDE.md from the published package
      • headless_ida_9.4_claude_skill
      • ida-domain
        • 1acb26e2: Add license control methods (#128)
      • ida-hcli
        • 13b42b24: fix(plugin bundle): download bundle wheels with uv, without IDA
        • 6220168a: test: restore a space dropped from an unrelated comment
        • 46477b24: refactor(plugin search): take the install hint's host from the query
        • 02ed142f: fix(plugin search): qualify the install hint for a colliding plugin name
        • b5a09e06: refactor(plugin lint): pass lint result explicitly instead of a globa…
        • 295358f1: fix(plugin lint): exit nonzero when lint reports errors
        • 2a6aedab: docs: reword ida-install-dir blank-value comment
        • dea32066: fix: reject empty and relative ida-install-dir values (#400)
        • e25e6b68: fix: run a Linux installer the user cannot make executable (#397)
        • 833ca0dc: fix(license install): handle a missing IDA_DIR without a terminal
        • 1e915bed: fix: report a missing IDA installation or idat in python doctor and e…
        • af4d7f44: fix(ida install): exit 1 when the install dir exists without IDA
        • 402dda1c: remove docstring note about -debugtrace
        • cb8b6ce3: fix(ida install): do not write the installer debug trace to the curre…
        • da228633: docs: show the output hcli plugin lint actually prints (#390)
        • aad32dbb: docs: list all six .plugin.platforms values
        • 2a71a6ec: docs: list all six platforms that bundle create -platform all selects
        • 20e9ec77: fix(ida install): make -dry-run actions match the real install (#399)
      • idamcp
        • a0458798: Report the type of idapython_eval's result
        • f7137458: Seed idapython_eval namespaces with ida_domain
        • a5399fe7: Await non-coroutine results of idapython_eval
        • bcf7cbe5: Mark the GUI server as started only after its thread starts
        • 65ff3fd8: Free a crashed instance's quota slot when it is reopened
        • f3a13118: Merge pull request #26 from doomedraven/p21-lock-liveness
        • bdbffa2f: Don't kill a headless instance right after starting it
      • oh-my-pi
        • f84deef3: docs(fork): list patch 35 in AGENTS.md
        • 44438c52: docs(fork): record patch 35 (lightweight mode)
        • 938ed7e9: feat(cli): -lightweight zero-injection knowledge Q&A mode
        • aed59ebd: docs(fork): record the v18.6.0 and v18.8.0 upgrades
        • 4d0dcfa3: upgrade/v18.8.0: merge upstream v18.8.0 (via trial/v18.8.0)
        • d47da844: test(ai): accept fork block-form user turns in compaction replay and …
        • 0fa1ff11: fix(catalog): exempt fork patch compat fields from parity and align g…
        • 1f9c0ca7: fix(ai): guard compat.extraBetas union against rows without the key
        • ffe3067e: Merge upstream v18.8.0 into the fork (trial)
        • 8d9f3a63: v18.8.0 upstream source
      • rhabdomancer
        • 772cc7a4: chore: improve Cargo.toml before publishing
        • 2bddae93: chore: prepare for release
        • 5c157ed4: style: improve lookup
        • 583c78bb: style: switch to a FoundFunction struct with an enum to keep the fu…
    2. 🔗 backnotprop/plannotator v0.28.8 release

      Follow @plannotator on X for updates

      Missed recent releases? Release | Highlights
      ---|---
      v0.28.7 | Ask this session keeps your drafts as drafts, image comments survive a reload, hostname check on every server
      v0.28.6 | Ask AI names the lines you selected, reorder quick labels and edit their emoji, Pi fixed-port crash fixed, OpenCode 2 subagent notice fixed
      v0.28.5 | Several files in one review, the plannotator tool on Pi and OpenCode 2, decisions name the exact file, Ask this session reconnects after sleep
      v0.28.4 | Comments come back after the agent edits a file, the agent can list and close its reviews, PR-description feedback no longer dropped, Pi thinking levels
      v0.28.3 | Typing into your agent while it answers an Ask no longer streams into Plannotator
      v0.28.2 | Question cards show wrapped choices and tables, pinned images named by their file, Done with nothing to send starts no agent turn
      v0.28.1 | Ask AI from diagram comments, pinned images named for the agent, OpenCode 1 URL toasts and one reply per feedback, folder feedback sent once
      v0.28.0 | Ask this session in Claude Code, Pi and OpenCode 2, Claude Code mod on by default, Pi plan review no longer blocks, first-run demo
      v0.27.25 | Code review works with color.diff = always, Bitbucket review fixes, wide tables no longer collapse in Firefox, install script fix
      v0.27.24 | Image previews stay in the all-files view, PR comment previews open on the commented line
      v0.27.23 | Bitbucket Cloud PR review, Question UI for answering agents in place, opt-in auto-update, viewed files remembered, review another repo or worktree

      What's New in v0.28.8

      This release brings the Plannotator Inbox (preview): one local window where your agents leave you messages, questions, files and guided reviews, and where your reply goes back to the session that asked.

      The Plannotator Inbox (preview)

      Agents often need you for one thing: a choice, a check, a look at a file. Until now that meant a review tab per question, or a question buried in the terminal. The Inbox collects all of it in one window on your machine:

      • Threads by project. Each agent conversation is a row, grouped in sections: Stopped on you, Holding up work, Waiting on you, Sent, New since you looked, and Quiet. Filter by project in the sidebar.
      • Questions you answer with a click. Agents ask with the same :::question cards plan review uses. Pick, add a note, press Send.
      • Files and annotations. Files an agent attached open beside the thread (markdown, text, HTML, Mermaid, Graphviz). Annotate them as in Plannotator; the annotations go with your reply. If the agent changed a file after sending it, the Inbox says so and can show the version it sent.
      • Decisions. A question can record your answer as a project decision, listed on a Decisions page.
      • Guided reviews. An agent can send a guided review of a code change into a thread.
      • New message. Write to an agent session that is running now, without waiting for it to ask.
      • Browser notifications while the tab is in the background, if you allow them.

      Your reply wakes the agent that asked. Claude Code (with the Plannotator mod) gets a plannotator_inbox tool, on by default. Pi and OpenCode 2 get the same tool, off by default because they send every tool's definition with each request: turn it on with PLANNOTATOR_INBOX_TOOL=1. Any other agent can connect through MCP with plannotator inbox mcp; the Inbox's Settings show the exact command for your agent.

      It is local only: it listens on 127.0.0.1, needs no account, and keeps everything under ~/.plannotator/inbox. Start it with:

      plannotator inbox
      

      You never have to keep it running: an agent starts it in the background when it needs it, without opening a tab. plannotator uninstall --purge removes its data. Full reference: plannotator.ai/docs/reference/inbox.

      Additional Changes

      • Question cards for host apps. @plannotator/ui 0.52.1 can show the "Records a decision" switch on any question and open a host's own decision card from it. Plannotator's own cards are unchanged (#1753).

      Install / Update

      macOS / Linux:

      curl -fsSL https://plannotator.ai/install.sh | bash
      

      Windows:

      irm https://plannotator.ai/install.ps1 | iex
      

      Claude Code Plugin: The plugin and the plannotator binary update separately, so run the install script above as well. In a terminal:

      claude plugin marketplace update plannotator
      claude plugin update plannotator@plannotator
      

      Then restart Claude Code. Inside Claude Code, run /plugin marketplace update plannotator, then open /plugin → Installed → plannotator → Update now.

      Pi:

      pi update --extensions
      

      OpenCode: Re-run the install script above. It now also clears the OpenCode 2 plugin cache.

      What's Changed

      Full Changelog : v0.28.7...v0.28.8

    3. 🔗 earendil-works/pi v1.1.0 release

      New Features

      • Program status reporting : terminals and agent dashboards that support OSC 7501 see whether Pi is working, blocked on a dialog or login, done, or failed. See Program status.
      • Claude Haiku 5.5 : anthropic/claude-haiku-5-5, with adaptive thinking up to xhigh/max effort.
      • Adjust default tools with+name/-name: --tools entries like pi -t +codemode,-write change the default selection instead of replacing it. See Tools.
      • GPT-6 Luna and image classification : OpenAI's GPT-6 Luna is available as a classifier model through the Decisions API, and codemode's models.classify() accepts images for classifiers that support them. See Use classifier models.
      • Native llama.cpp decision models : Julia-1, Laya, Kev, lev, and OpenJev served by llama.cpp 0.6.0 or later run natively as classifiers through /v1/systemone. See Classification.

      Added

      • Added +name and -name entries to --tools, which change the default tool selection instead of replacing it, for example pi -t +codemode
      • Added durationMs to the tool render context and to tool_execution_end extension events: the recorded execution time of a final tool result (#10549)
      • Added outputPad to the tool render context (#10557 by @rwachtler)
      • Added OpenAI's GPT-6 Luna as a classifier model through the Decisions API, available with OPENAI_API_KEY (see Use classifier models)
      • Added images to codemode's models.classify() context, so classifiers that accept images, such as GPT-6 Luna, can judge them
      • Added program status reporting with OSC 7501: terminals and agent dashboards that support it see whether Pi is working, blocked on a dialog or login, done, or failed. PI_PROGRAM_STATUS=1|0 overrides detection (see Terminal setup) (#10607)
      • Added aborted to agent_settled session, extension, and JSON events, so integrations can tell a cancelled run from a finished one (#10607)
      • Added Claude Haiku 5.5 (anthropic/claude-haiku-5-5), with adaptive thinking up to xhigh/max effort and prompt caching on Bedrock
      • Added native llama.cpp decision models: Julia-1, Laya, Kev, lev, and OpenJev served by llama.cpp 0.6.0 or later are listed only as classifiers through /v1/systemone instead of as chat models (see Classification) (#10382)

      Changed

      • Changed outputPad to also apply to ! command output, tool output, and summary blocks (#9946, #10557 by @rwachtler)
      • Changed pi mcp login --timeout to limit the whole sign-in, including requests to the authorization server, instead of only the wait for the browser (#10565)

      Fixed

      • Fixed bash and PowerShell results losing Took after reloading a session, and the live Took including wall-clock steps; both now show the recorded execution time (#10549)
      • Fixed managed installs keeping every old release; pi update now keeps only the new release and the one it updated from (#10392, #10511 by @davidbrai)
      • Fixed standalone binaries loading .env, .env.local, and .env.development from the launch directory into Pi's environment (#10473)
      • Fixed !! command headers losing their dim color once output arrives (#10557 by @rwachtler)
      • Fixed the codemode description not marking searchTools(), describeTool(), and describeNamespace() as async, which led models to serialize the unawaited promise as {} (#10555)
      • Fixed codemode output items running together, so models could not tell where one text() or console.log() output ended and the next began. With several text items, each now starts with a ==> text N/M <== line, and console calls follow the other output in one <console_output> block with one line per call
      • Fixed /mcp waiting for all servers to connect before opening; the manager now updates live and remains usable while enabling, reconnecting, or disabling servers (#10562)
      • Fixed images being dropped as "could not be resized" when running under node --watch on Node 24.19+ and 26.x, where Node posts its own messages on the image resize worker channel (#10527)
      • Fixed clipboard paste doing nothing in Termux, and failed copies there omitting the Termux:API install hint (#10391)
      • Fixed ! and RPC bash output keeping fragments of color codes, such as a stray m, when a code was split across output chunks (#10504)
      • Fixed MCP OAuth sign-ins that could not be cancelled while waiting on the authorization server and kept running after the session ended. The sign-in screen now cancels with Esc at every step, session shutdown aborts a running sign-in, and each request to the authorization server times out after 15 seconds (#10565)
      • Fixed shutdown waiting up to 15 seconds to refresh an MCP OAuth token that was about to expire, only to close the server's session (#10565)
      • Fixed the fullscreen text selection surviving session switches and other transcript rebuilds, which highlighted unrelated text in the new transcript (#9311, #10567 by @christianklotz)
      • Fixed OpenAI models on Bedrock ignoring the thinking level and always running at Bedrock's default reasoning effort (#9331, #10142 by @jsanter27)
      • Fixed model headers in models.json not overriding the originator and User-Agent headers of Codex requests (#10429 by @lucasmeijer)
      • Fixed server_busy and servers are currently busy provider errors ending the turn instead of being retried (#10543)
      • Fixed Mistral responses that end with finish_reason: "error" not being retried (#10487)
      • Reduced context-limit request failures by estimating input at 3.5 characters per token instead of 4 when calculating output limits (#10497)
      • Fixed Radius models disabled by an organization owner still being listed
      • Fixed Anthropic browser login failing with "localhost refused to connect" when port 53692 is reserved or in use, for example by Hyper-V/WSL port exclusions on Windows: login now falls back to a free loopback port (#10571)
      • Fixed session costs undercounting long prompts on models with prompt-length pricing tiers, such as Claude Haiku 5.5, Gemini 3.1 Pro, and GPT-5.4, through OpenCode, OpenCode Go, OpenRouter, Vercel AI Gateway, Google, MiniMax, and other providers
      • Fixed Markdown links not being clickable in Herdr (#10573)
    4. 🔗 jj-vcs/jj v0.46.0 release

      About

      jj is a Git-compatible version control system that is both simple and powerful. See
      the installation instructions to get started.

      Release highlights

      • Jujutsu can now colocate workspaces besides the default one by creating Git
        worktrees. Use jj workspace add --[no-]colocate and the setting
        git.colocate to control this.

      Breaking changes

      • The minimum supported git command version is now 2.42.0, up from 2.41.0.
        jj workspace add uses git worktree add --orphan, which was added in
        2.42.0.

      • The minimum supported Rust version (MSRV) is now 1.97.1.

      • jj bisect run now runs some consistency checks before proceeding to bisect.
        This helps ensure that the command can tell good and bad revisions apart,
        and that the working copy does go from bad to good over the provided revset.
        Use the new flag --trust-endpoints to disable these checks.

      • jj split now opens a single editor session to edit descriptions for the
        split commits.

      • jj undo and jj redo now refuse to undo/redo an operation that was
        performed in another workspace. Use --allow-cross-workspace to undo/redo
        it anyway.

      • jj workspace list/root no longer omit unreachable paths. All recorded
        paths are now shown, with warnings displayed in jj workspace root.

      • The List.get(), .first(), and .last() template functions now return
        Option<T> instead of throwing an error on out-of-bounds access.

      New features

      • jj workspace add supports --colocate/--no-colocate flags to control
        whether a Git worktree is created alongside the workspace. The default
        colocates when the current workspace is colocated and the git.colocate
        config is true. jj workspace forget removes the corresponding Git
        worktree when one exists.

      • jj git colocation status/enable/disable now work on child
        workspaces. status correctly reports colocation state and includes
        the workspace name. enable creates a Git worktree and disable
        removes it, allowing colocation to be toggled after workspace
        creation.

      • jj workspace remove removes a workspace and its directory from disk. The
        working-copy state is snapshotted into a commit before removal.

      • Added commands jj file edit and jj file delete for editing files in any
        revision without needing to change the working copy.

      • jj git push now supports pushing to multiple remotes at the same time.
        This can be configured via git.push set to a string pattern
        or array of string patterns, or with the repeatable --remote flag,
        which also accepts string patterns.

      • The default target revisions for jj git push can now be configured via
        revsets.git-push.

      • Added the TreeEntry.normal_value() template method and the TreeValue type
        to access resolved tree values, formatted as their full object IDs, including
        Git submodule commit IDs.

      • Diff hunk headers now include nearby source symbols for many common
        programming and markup languages.

      • fix.tools.<name>.line-range-args (replaces line-range-arg) is an array of
        string template args to pass to the fix tool. This is more flexible in cases
        where you need to pass multiple arguments to the tool, such as separate args
        for the range start and range end.

      • jj run now uses the sparse patterns from the workspace it's run from.
        Use the --sparse-patterns option to control this behavior (evaluated
        per each jj run invocation).

      • jj util diff <path1> <path2> to compare files on disk.

      • Aliases now support setting aliases.<name>.enabled = false, which will
        disable them. This can be used to disable built-in aliases or disable aliases
        in later layers (such as repo config files).

      • ui.editor now supports $path and $line substitution variables. Example:
        ui.editor = ["emacs", "+$line", "$path"]

      • fill template function now supports an additional named parameter
        break_words, that allows specifying if the template should break words
        longer than width passed in the input to ensure no words overflow the
        specified width.

      • The json() template function now supports map literals: json({'key' => value})

      • The hunk headers of diff.color-words.conflict = "pair" now include the
        conflict labels of the compared terms.

      Fixed bugs

      • On Windows, jj no longer hangs when a subprocess needs to prompt the user,
        such as ssh asking for a key passphrase or for confirmation of an unknown
        host key. Subprocesses started from a terminal now inherit its console, rather
        than being given an invisible one by CREATE_NO_WINDOW for the prompt to
        disappear into.
        #6745
        #8547

      • On Windows, jj git colocation enable and jj git colocation disable no
        longer fail with "Access is denied (os error 5)" when the Git repository
        contains pack files.
        #8661

      • jj undo of jj workspace forget now correctly preserves the workspace's
        recorded path. Previously the path metadata was lost, leaving the workspace
        in a broken state after undo.
        #9991

      • at_operation() can now be used with operations that are not ancestors of
        the current operation (e.g. sibling operations created by concurrent
        commands). Previously, evaluating such expressions failed if they resolved
        to commits missing from the current operation's index.

      • .gitignore files are now respected even if they aren't materialized in the
        working copy because they are excluded by the sparse patterns. Previously,
        ignored files could become tracked in a sparse working copy.
        #2289

      • In-tree ignore files (.gitignore) are no longer read through symlinks,
        matching git behavior. Such files are now silently skipped instead of having
        their symlink target applied. $GIT_DIR/info/exclude and core.excludesFile
        are unaffected and still follow symlinks, as git does.
        #7161

      • jj workspace list templates are now labeled with workspace name,
        workspace root, etc.

      Contributors

      Thanks to the people who made this release happen!

    5. 🔗 Simon Willison Claude Haiku 5.5 rss

      As previously promised, here's Anthropic's new fast, low cost model: Introducing Claude Haiku 5.5.

      The previous Haiku, 4.5, was very much showing its age. It came out almost a year ago, and was priced at $1/million input and $5/million output - relatively expensive even back then, and a full 10x the price of OpenAI's GPT-6 Luna, released last month.

      The new Haiku exactly matches the price of GPT-6 Luna - $0.10/$0.50 - up to 100,000 tokens. Beyond 100,000 tokens the price increases 5x to $0.50/$2.50. Luna itself has a price increase at 272,000 tokens but only to $0.20/$0.75.

      Haiku 5.5 also uses a new, less generous tokenizer. My Claude Token Counter tool shows that the same long prompt uses around 1.25x as many tokens with Haiku 5.5 compared to Haiku 4.5, so there's a hidden price increase there.

      If your workloads fit in 100,000 tokens, Haiku is the same price as Luna and reports higher benchmark scores. Above 100,000 tokens, Luna looks like a much better deal.

      The most recent release of llm-anthropic finally fixed it so I don't need to ship a new version of that plugin for every new model. I tested the new model like this:

      llm install -U llm-anthropic
      llm anthropic refresh
      llm -m claude-haiku-5.5 "Generate an SVG of a pelican riding a bicycle" -o thinking_effort low
      

      Pelicans

      Here are pelicans for low, medium, high, xhigh, and max. The new Haiku doesn't let you disable reasoning, and defaults to medium. I got a good bicycle frame for everything beyond low. The low effort pelican cost 0.0936 cents and took 7 seconds.

      This max effort pelican (with a reasoning trace that starts "This is the classic pelican-on-bicycle SVG test...") took 5 minutes 9 seconds to generate, but still only cost me 3.3826 cents:

      Flat vector illustration: a white pelican wearing a red cap rides a dark grey bicycle from left to right across a green grassy strip, its long orange legs reaching down to the orange pedals, a grey wing stretched forward to grip the curved handlebars, and its large yellow-orange beak with a bulging orange throat pouch pointing ahead; behind it, white speed lines show motion, and above, a yellow sun with a pale halo and two white clouds sit in a gradient blue sky.

      (Since the reasoning trace exhibits awareness of the benchmark, here's Generate an SVG of an armadillo in fishnet tights jaywalking on Mars (on xhigh), and the same prompt against some other recent models. Background on that.)

      For comparison, here's the pelican I got a year ago from Haiku 4.5 (for 0.7583 cents - Haiku 4.5 did not support reasoning levels). It sucked at drawing pelicans:

      Described by Haiku 4.5: A whimsical illustration of a bird with a round tan body, pink beak, and orange legs riding a bicycle against a blue sky and green grass background.

      And a generous API credit scheme for subscribers

      In addition to Haiku 5.5, Anthropic announced today that they are halving the price of cache reads for Sonnet 5.5. They've also added API credits to subscription plans:

      Second, this week, we’ll roll out a new monthly API credit to all Max and Team subscribers for use on the Claude Platform. Max 5x users will get $100 in credits per month, Max 20x users will get $200, and Team subscribers will receive up to $500, pooled across their users.

      Claiming this is pleasantly easy: navigate to Settings -> Billing and select the API organization that should benefit from the credits every month:

      Screenshot of a Settings dialog with a close X button at top right and a row of tabs reading "General", "Account", "Privacy", "Billing" (selected), "Usage", "Capabilities", "Memory" and "Design sys" (cut off at the edge). Beside a line-drawn icon of a branching plant with circular buds, the panel reads "Max plan", "20x more usage than Pro" and "Your subscription will auto renew on Nov 2, 2026." with an "Adjust plan" button on the right. Below is a section headed "API credits" with a blue "New" badge, reading "Your Max 20x plan includes $200 USD in API credits each month." followed by grey text "Link an API organization to start receiving credits." with a "Link organization" button on the right.

      The API credits exactly match the cost of the subscription itself. This is really generous - it makes it much easier for subscribers to use the API. Anthropic also let you disable auto-reload for the API, with the consequence that "API requests will stop when your balance runs out" - exactly what you want if you're planning to burn through those API credits without risk of a nasty billing surprise.

      Note that the monthly credits do not roll over - use them or lose them.

      OpenAI still allow you to use your Codex subscription for personal API use, which works out as a better deal for heavy API users. This new credit scheme goes at least some way to overcoming that difference.

      You are only seeing the long-form articles from my blog. Subscribe to /atom/everything/ to get all of my posts, or take a look at my other subscription options.

    6. 🔗 r/Harrogate New to Harrogate and wondering about traffic around the centre rss

      Due to getting a new job in Leeds and moving out of a relatively leafy part of London, I’m thinking of a move to Harrogate and looking at places not too far from the town centre. I’ve been looking at places on Ripon Road (and the roads just off it) near(ish) to the Doubletree Hilton. Every time I look at Google Maps, it seems to show heavy , often stationary, traffic in both directions and particularly around Crescent Road. Would living around there be an eternal traffic nightmare, or just at peak times of the working week?

      submitted by /u/Creepy_Ad6865
      [link] [comments]

    7. 🔗 exe.dev Crossing the Hyper-Thread Boundary rss

      Modern processors have multiple CPU cores, and the physical CPU cores in turn have two logical CPU cores. The latter permit the physical core to run two independent programs simultaneously, an approach known as hyper- threading. Hyper-threading uses CPU resources more efficiently, but it exposes transient execution vulnerabilities, in which a program running on one hyper-thread is able to extract data from the program running on the other hyper-thread. An example of such an attack is MDS.

      exe.dev runs virtual machines on behalf of different users. We need to protect against the possibility of one user exploiting a vulnerability to extract data from another user running on a different hyper- thread of the same CPU core. Fortunately the Linux kernel supports core scheduling cookies to control which processes are permitted to share a CPU core.

      It is straightforward to give each VM an independent cookie, meaning that two different VMs never run on the same CPU core. However, in practice this leads to measurably inefficient use of physical CPU cores. So we instead implemented a more efficient, but still secure, mechanism: the VMs of each team use an independent cookie. This means that two different VMs from a different team (or from a different user for users not on teams) never run on the same CPU core. To put it another way, we assume that different members of the same team trust each other.

      Of course, for various reasons, some teams do not have trust among all their VMs. An admin of those teams can run ssh exe.dev team settings core-sharing off to prevent their VMs from sharing CPU cores. (Currently VMs that do not belong to a team never share cores with other VMs, even VMs from the same user; if that changes someday we will introduce a similar setting for individuals.)

    8. 🔗 backnotprop/plannotator v0.28.7 release

      Follow @plannotator on X for updates

      Missed recent releases? Release | Highlights
      ---|---
      v0.28.6 | Ask AI names the lines you selected, reorder quick labels and edit their emoji, Pi fixed-port crash fixed, OpenCode 2 subagent notice fixed
      v0.28.5 | Several files in one review, the plannotator tool on Pi and OpenCode 2, decisions name the exact file, Ask this session reconnects after sleep
      v0.28.4 | Comments come back after the agent edits a file, the agent can list and close its reviews, PR-description feedback no longer dropped, Pi thinking levels
      v0.28.3 | Typing into your agent while it answers an Ask no longer streams into Plannotator
      v0.28.2 | Question cards show wrapped choices and tables, pinned images named by their file, Done with nothing to send starts no agent turn
      v0.28.1 | Ask AI from diagram comments, pinned images named for the agent, OpenCode 1 URL toasts and one reply per feedback, folder feedback sent once
      v0.28.0 | Ask this session in Claude Code, Pi and OpenCode 2, Claude Code mod on by default, Pi plan review no longer blocks, first-run demo
      v0.27.25 | Code review works with color.diff = always, Bitbucket review fixes, wide tables no longer collapse in Firefox, install script fix
      v0.27.24 | Image previews stay in the all-files view, PR comment previews open on the commented line
      v0.27.23 | Bitbucket Cloud PR review, Question UI for answering agents in place, opt-in auto-update, viewed files remembered, review another repo or worktree
      v0.27.22 | Plans open the docs they link to, Claude review jobs locked down, Code Tour on Linux without Claude's sandbox, Pi reviews the latest plan

      What's New in v0.28.7

      Seven PRs, two from community contributors, one of them a first-time contributor. The main fix stops Ask this session from handing your unsubmitted comments to the agent as instructions. Other changes: comments on images survive a reload, other plugins' commands no longer carry Plannotator's name, and every Plannotator server now checks the hostname it was reached by.

      Ask this session keeps your drafts as drafts

      When you asked a question through "Ask this session", Plannotator also sent the comments you had not submitted yet, in the same format as submitted feedback. The agent read them as instructions and could act on them before you pressed Submit. In the report, it created GitHub issues from draft comments. On Submit, the same comments then arrived a second time.

      Draft comments are now sent once, as a short list clearly marked as drafts you have not submitted, with an instruction not to act on them. Your question always comes last. The list goes again only when it changes. Submit delivers your feedback once, as before. This applies on Claude Code, Pi and OpenCode, and to the annotate agent terminal. The separate Ask AI providers are unchanged. (#1749, closing #1748, reported by @krizman)

      Comments on images come back after a reload

      In HTML annotate, a pin on an image, video or embedded page that had no id disappeared after a reload or an HTML refresh, which made image comments in generated reports hard to keep. Such pins are now found again by the file they show. A pin follows its image when other images are added around it. If the image's source changes, the pin shows as no longer matching. If the same source appears twice, the pin is not saved, so it can never land on the wrong copy. Live-app sessions behave the same way. (#1747 by @pro- vi)

      Other plugins' commands no longer carry Plannotator's name

      With the Plannotator mod on, Claude Code labelled the output of any other plugin's slash command as if Plannotator had helped answer it (for example plannotator+frontend: …). The mod now listens only to its own commands. /plannotator-review, /plannotator-annotate and /plannotator-last work as before, with or without the Plannotator skills installed. (#1741 by @rushelex, closing #1740)

      Every server checks the hostname it was reached by

      Plannotator's review servers now refuse requests that arrive under an unexpected hostname. A normal local session accepts only localhost and loopback addresses. Remote mode also accepts IP addresses, your configured PLANNOTATOR_URL_HOST and the machine's own name (including name.local). --tailscale sessions accept their tailnet name, and code-server and Coder proxies are recognised from VSCODE_PROXY_URI. VS Code, SSH and Docker port forwarding, Codespaces and dev tunnels keep working as before.

      If you reach a remote-mode session by another DNS name (a server's public name, for example), you now get a 403 that says what to do: set PLANNOTATOR_URL_HOST, or list the name in the new PLANNOTATOR_ALLOWED_HOSTS (comma-separated; * turns the check off). (#1742)

      Additional Changes

      • OpenCode 2: extra words no longer stop annotate. /plannotator-annotate . notes.md (or please notes.md) opened nothing and showed nothing. It now opens the file, and a slash command that fails shows the reason in the session instead of only in OpenCode's log. On OpenCode 1 this applies when the plugin uses the CLI runtime; the default embedded runtime is unchanged (#1739).
      • Question cards for host apps. @plannotator/ui 0.52.0 lets an app embedding the question cards turn "Records a decision" on and off from the card's header and hide the card's own status tag. Plannotator's own cards are unchanged (#1744).

      Install / Update

      macOS / Linux:

      curl -fsSL https://plannotator.ai/install.sh | bash
      

      Windows:

      irm https://plannotator.ai/install.ps1 | iex
      

      Claude Code Plugin: The plugin and the plannotator binary update separately, so run the install script above as well. In a terminal:

      claude plugin marketplace update plannotator
      claude plugin update plannotator@plannotator
      

      Then restart Claude Code. Inside Claude Code, run /plugin marketplace update plannotator, then open /plugin → Installed → plannotator → Update now.

      Pi:

      pi update --extensions
      

      OpenCode: Re-run the install script above. It now also clears the OpenCode 2 plugin cache.

      What's Changed

      • Ask this session: unsubmitted annotations are never sent as feedback by @backnotprop in #1749
      • fix(annotate): pins on images without an id restore by their source by @pro-vi in #1747
      • fix(mod): match command.run by command name by @rushelex in #1741
      • Validate the Host header on every request (defense in depth) by @backnotprop in #1742
      • OpenCode 2: /plannotator-annotate with extra words opens the file, and failures are shown by @backnotprop in #1739
      • ui: host-controlled decision toggle and status tag on question cards by @backnotprop in #1744
      • Groundwork for a feature that is not released yet, hidden from help and the agent skill, by @backnotprop in #1745 and #1750

      New Contributors

      Contributors

      @pro-vi found that comments on images without an id were lost on every reload, and wrote the fix that finds them again by their source. @rushelex reported the +plannotator label on other plugins' commands and fixed it with a one-line matcher. It is their second contribution, after the VS Code clipboard fix in #970.

      Community:

      • @krizman reported the Ask this session draft problem with exact steps, the text the agent received, and the likely cause (#1748).

      Full Changelog : v0.28.6...v0.28.7

    9. 🔗 r/Harrogate Fresh Stop - what’s going on rss

      Has anyone been to the shop near Asda? It’s one of those little grocery/corner-shop type places where you’d normally expect them to sell Pepsi, Coke, vapes, snacks, stuff like that.

      I went in a few times now there was basically nothing on the shelves. It was really odd. Does anyone know what the deal is with it? Is it some kind of front/drug thing, or is there a completely normal explanation I’m missing?

      submitted by /u/FrontRaspberry5060
      [link] [comments]

    10. 🔗 backnotprop/plannotator v0.28.6 release

      Follow @plannotator on X for updates

      Missed recent releases? Release | Highlights
      ---|---
      v0.28.5 | Several files in one review, the plannotator tool on Pi and OpenCode 2, decisions name the exact file, Ask this session reconnects after sleep
      v0.28.4 | Comments come back after the agent edits a file, the agent can list and close its reviews, PR-description feedback no longer dropped, Pi thinking levels
      v0.28.3 | Typing into your agent while it answers an Ask no longer streams into Plannotator
      v0.28.2 | Question cards show wrapped choices and tables, pinned images named by their file, Done with nothing to send starts no agent turn
      v0.28.1 | Ask AI from diagram comments, pinned images named for the agent, OpenCode 1 URL toasts and one reply per feedback, folder feedback sent once
      v0.28.0 | Ask this session in Claude Code, Pi and OpenCode 2, Claude Code mod on by default, Pi plan review no longer blocks, first-run demo
      v0.27.25 | Code review works with color.diff = always, Bitbucket review fixes, wide tables no longer collapse in Firefox, install script fix
      v0.27.24 | Image previews stay in the all-files view, PR comment previews open on the commented line
      v0.27.23 | Bitbucket Cloud PR review, Question UI for answering agents in place, opt-in auto-update, viewed files remembered, review another repo or worktree
      v0.27.22 | Plans open the docs they link to, Claude review jobs locked down, Code Tour on Linux without Claude's sandbox, Pi reviews the latest plan
      v0.27.21 | Remote and phone sessions load several times faster, real Request changes on GitHub, model pickers show real names, OpenCode fixes

      What's New in v0.28.6

      Six PRs, two of them asked for by the community. This is a fix release: Ask AI tells the agent which lines you selected, quick labels can be reordered and given a new emoji, and a few problems found while testing 0.28.5 on Pi and OpenCode are fixed.

      Ask AI says which lines you selected

      When you select text in a markdown document and ask a question, the question now names the lines, for example Source: /path/to/doc.md, line 41, or lines 41–44 when the selection spans several lines or blocks. Before, it carried only the file path, so when the same phrase appeared twice the agent could not tell which one you meant. The line numbers are the same ones the exported comment for that selection prints, so the question and your feedback point at the same place. This works in plan review and annotate, including linked documents and folder sessions, and through the side chat, Ask this session and the annotate agent terminal. (#1732, closing #1731, requested by @de-tre)

      Reorder quick labels and change their emoji

      Settings → Labels now has Move up and Move down buttons on each quick label, and the emoji is an editable field. The list order decides which Alt/⌥ number applies a label, and the key hint on every row updates as you move things. The emoji field takes exactly one emoji (flags, skin tones and combined emoji included); anything else is shown as invalid and never saved. Two labels with the same text no longer get mixed up when you edit one of them. Your saved labels keep the same format, so nothing needs migrating. (#1738, closing #1736, requested by @RobertoArtiles)

      Pi no longer crashes when a fixed port is busy

      With PLANNOTATOR_PORT set in a local Pi session, opening a second /plannotator-annotate while one was already open killed Pi with EADDRINUSE. The new review now takes the port over from the old one, as remote mode already did, and Pi stays up. (#1733)

      OpenCode 2: a review opened by a background subagent no longer adds a

      stray turn

      With the plannotator tool turned on, a background subagent that opened a review posted its "session ready" link into the main session while it was idle. The next time the main session woke up, the model answered that link line as its own turn, and the exchange stayed in every later request. The link now goes to the session that called the tool, while that call is still open, so it never becomes a turn of its own. Decisions were always delivered correctly and still go to the main session. (#1734)

      Additional Changes

      • Review names match what opened. On Claude Code, /plannotator-annotate . a.md opens only a.md, but the mod named the review "2 files: ., a.md" in its reply, status line and decision. Reviews are now named from what the CLI actually opened, on Claude Code and for OpenCode 2's tool (#1737).
      • Contributors' OpenCode setup left alone. Running bun install in a Plannotator checkout overwrote the developer's own OpenCode commands and skill. The OpenCode plugin's install step now copies them only when the package is installed under node_modules, and it works on Windows. Set PLANNOTATOR_OPENCODE_POSTINSTALL=1 or 0 to force it either way (#1735).

      Install / Update

      macOS / Linux:

      curl -fsSL https://plannotator.ai/install.sh | bash
      

      Windows:

      irm https://plannotator.ai/install.ps1 | iex
      

      Claude Code Plugin: The plugin and the plannotator binary update separately, so run the install script above as well. In a terminal:

      claude plugin marketplace update plannotator
      claude plugin update plannotator@plannotator
      

      Then restart Claude Code. Inside Claude Code, run /plugin marketplace update plannotator, then open /plugin → Installed → plannotator → Update now.

      Pi:

      pi update --extensions
      

      OpenCode: Re-run the install script above. It now also clears the OpenCode 2 plugin cache.

      What's Changed

      • Ask AI: name a text selection's source lines in the question by @backnotprop in #1732
      • Settings → Labels: reorder quick labels and edit their emoji by @backnotprop in #1738
      • fix(pi): attach the annotate agent terminal after listen so a busy fixed port cannot crash Pi by @backnotprop in #1733
      • fix(opencode): a tool launch's session-URL notice never wakes an idle root by @backnotprop in #1734
      • fix(mod, opencode): name a review by what the CLI opened, not the typed words by @backnotprop in #1737
      • fix(opencode): skip the plugin postinstall inside the monorepo (cross-platform) by @backnotprop in #1735

      Community

      • @de-tre asked for the selected lines to be included in Ask AI questions, so the agent knows which occurrence of a phrase they meant (#1731).
      • @RobertoArtiles asked for a way to reorder quick labels and change their emoji (#1736).

      Full Changelog : v0.28.5...v0.28.6

    11. 🔗 matklad On Git Refs rss

      On Git Refs

      Oct 7, 2026

      I have recently improved my mental model of Git. Consider these two git commands:

      $ git fetch origin master
      $ git switch -c my-feature origin/master
      

      Do you understand why is it origin master in one command, and origin/master in the other? I didn’t, until a few weeks ago!

      My understanding was that git is a content-addressable database. Git stores commits, a commit is identified by the hash of its content, and the content of a commit is, primarily:

      • a memory-less snapshot of a state of the codebase at a given point in time,
      • a list of (hashes of) parent commits.

      That was enough git for me to understand git log output and get me out of any botched rebase without having to re-clone the repo (For roughly half of my career, I was re-cloning the repo. No shame in that! Learning git is useful, but it’s not the highest priority thing to learn when you start).


      I now understand that git not only comes with an append-only (“immutable”) content-addressable database, but is also a boring mutable key-value store.

      Git has a mutable map whose keys are strings, and whose values are content- addressed objects. The keys are conventionally formatted as file system paths, and you can usually inspect the state of the mapping by listing .git/refs directory:

      $ eza -T .git/refs
      .git/refs
      ├── heads
      │   ├── make
      │   ├── master
      │   ├── my-feature
      │   └── pbd-adt
      ├── origin
      ├── remotes
      │   └── origin
      │       ├── context-switches
      │       ├── gh-pages
      │       ├── HEAD
      │       ├── make
      │       └── master
      └── tags
      
      $ cat .git/refs/heads/master
      b59148228e52f7c615ead7fdd4e91001994ad50f
      
      $ git show-ref refs/heads/master
      b59148228e52f7c615ead7fdd4e91001994ad50f refs/heads/master
      

      What makes this refs KV infrastructure confusing is that:

      • It powers many distinct user-visible git features, but refs themselves are an implementation detail.
      • It is a bit of a leaky abstraction, refs are almost invisible in the day-to-day usage.
      • Git CLI uses shorthand notation for refs and many default arguments, which makes it not obvious that a particular CLI argument is a ref.
      • And, as usual, git likes to give several names to one thing, and re-uses the same name for distinct things.

      Branches, tags, and git notes are all just refs!


      The structure becomes much more obvious once you elaborate all CLI shortcuts. The original command

      $ git fetch origin master
      

      then becomes

      $ git fetch \
          https://github.com/matklad/matklad.github.io \
          refs/heads/master:refs/remotes/origin/master
      

      The first argument of fetch (https://...) is a location of a remote repository. Git will “dial” that address, and will transfer some data from that computer locally over the network.

      The second argument is a source:target pair of string keys (refs). The source is a key on the remote repo, the target is the name of a local key, and fetch as a whole asks git to read a value from a remote repository and save it locally under a different name.

      To avoid typing repository URLs all the time, git assigns them symbolic names, with origin being the conventional name for the primary remote repository:

      $ git fetch origin \
          refs/heads/master:refs/remotes/origin/master
      

      refs/heads/master is a fully elaborated name of a branch on the remote repo. That is, branch my-feature is just a refs/heads/my-feature ref. It could have been refs/branch/my-feature, but it isn’t :)

      I don’t know the specific shorthand rules, but, generally, git allows you to spell only the suffix of a ref:

      $ git fetch origin \
          master:refs/remotes/origin/master
      

      refs/remotes/origin/master is the name of the local ref we’ll use to store the result. It would seem natural to just use the same name locally as the one on the remote, but this only works if there’s a single remote. If there are two upstream repositories (for example, your fork, and the original repo you forked from), their ref names will collide. That’s why we want to namespace the refs for remote called foo under refs/remotes/foo. And origin is just a conventional name for the remote in simple setups.

      Again, it would be more natural to directly mirror remote ref structure locally:

      refs/ heads/my-branch -> refs/ remotes/origin/ heads/my-branch
      

      but git strips the redundant heads component. And this -heads, +remotes/$remote mapping is built in, which compresses the command to

      $ git fetch origin master
      

      It’s worth reflecting why it works this way. Git model is offline first. What’s more, it assumes explicit synchronization points. Rather than synchronizing with the remote repository in background when there’s connectivity, git requires explicit fetch and push operations to transfer bytes over the wire. In this paradigm, it is useful to model the state of the remote party at the moment when we talked to them the last time. Theory of mind!

      This hopefully deconfuses git’s concept of local and remote branches. Consider the main branch. It exists on the remote named origin as refs/heads/main. When you synchronize your local repository with origin, you get refs/remotes/origin/main — you current best knowledge about the the state of main on the origin.

      And then there’s your local refs/heads/main. It typically starts pointing at the same commit as refs/remotes/origin/main. But, when you make a commit, refs/heads/main advances, but refs/remotes/origin/main stays the same.

      When you try to push your local commit to origin, you will get a conflict, if the main branch on the origin advanced in the meanwhile. In that case, git automatically updates refs/remotes/origin/main (as that’s just a local mirror of the remote state), but then it’s on you to update refs/heads/main and push it again.

      Revisiting the full example:

      $ git fetch origin master
      $ git switch -c my-feature origin/master
      

      The first command looks up the URL for the origin remote in .git/config and makes a network request to that machine. As a result, the local refs/remotes/origin/master gets updated to the same commit as refs/heads/master remotely (the commit and its ancestors are transferred locally as a result).

      The second command creates a refs/heads/my-feature ref (a branch), whose starting point is refs/remotes/origin/master. It is an example of a leaky abstraction.

      The second argument there is a (shorthand of a) ref, so you can do

      $ git switch -c my-feature \
          refs/remotes/origin/master
      

      But, although the first argument creates a ref, it isn’t a ref itself. In other words, if you try to elaborate it as well

      $ git switch -c refs/heads/my-feature \
          refs/remotes/origin/master
      

      you’ll get

      refs/heads/refs/heads/my-feature
      

      That’s all! I am pretty sure this isn’t particularly useful, but maybe it is interesting!

    12. 🔗 New Music Releases Philip Glass - Philip Glass: Music for Film rss

      Philip Glass - a new release is available:

      • 2026-10-07: Philip Glass: Music for Film (Album)

      Amazon: Canada | Deutschland | France | United Kingdom | United States

      Visit muspy for more information.

  4. October 06, 2026
    1. 🔗 IDA Plugin Updates IDA Plugin Updates on 2026-10-06 rss

      IDA Plugin Updates on 2026-10-06

      Activity:

      • capa
        • d4cf889f: Merge pull request #3175 from mandiant/dependabot/npm_and_yarn/web/ex…
        • f890a430: Merge pull request #3173 from mandiant/dependabot/npm_and_yarn/web/ex…
      • disrobe
        • f6fae606: test(dotnet): consolidate integration target
        • d52b6b70: test(py-decompile): consolidate integration target
        • 7743c9f6: test(jvm): consolidate integration target
        • 1f3c2e06: fix(xtask): simplify consolidated audit paths
        • 9b22833a: fix(xtask): match consolidated module filters
        • 8cf5b4e9: fix(xtask): scope explicit test module attributes
        • d542420b: fix(xtask): audit consolidated test modules
        • b1767a0c: fix(ci): quote consolidated test filters
        • 42abf9ec: test(cli): consolidate integration target
        • 5e978218: test(native): consolidate integration target
        • e7d81657: chore(xtask): regenerate the published figures
        • cc3e426e: test(js-deob): consolidate integration target
        • 13ee9ce6: test(testkit): load bounded runtime fixtures
      • distro
        • 41125aba: Build 14 PPA packages for arm64
      • mcrit-plugin
        • 3e3f6f2a: Merge pull request #42 from familiary/docs/hcli-1x-upgrade
        • cbad5092: Tell 1.x users to reinstall through HCLI after the transfer
      • rhabdomancer
        • 3a4dff38: doc: improve readme and changelog before release
        • 58e9a077: doc: futher improvements
        • 04dacdf1: doc: minor improvements
        • 8910d093: feat: label call locations in library code recognized by IDA with `(l…
        • a1f8bb58: feat: label call locations in library code recognized by IDA with `(l…
        • 2d778674: doc: add todo item for following calls through thunks outside .plt …
        • cd7ad31c: fix: stop marking instructions that fall through into a bad function,…
        • 75545f8b: fix: match the glibc aliases that IDA may pick over the plain names i…
        • f6b383e6: refactor: stub concept unification
        • 1c4146c4: doc: add a note about fortified function handling
        • fc1b1865: fix: match import stubs that IDA names with a numeric suffix and Univ…
      • Signature-Locator
        • 01539af9: feat: support manual input without length restriction and multi-langu…
      • symbolicator
        • c2ed4ef5: chore: refresh kernel signatures with IDA 9. + latest 27.2 beta KDK
    2. 🔗 Evan Schwartz Scour - September Update rss

      Hi friends,

      In September, Scour scoured 1.2 million articles (up from ~880,000 in August) from 28,568 feeds. Also, welcome to the 116 new users who signed up since my last product update email!

      Here's what's new in the product:

      📚 Library and Article Tabs

      You can now find all of your saved, loved, and liked posts, as well as your full reading history, in the Library section.

      Also, if you click Read on Scour for any article, that page now has tabs for the article's content, other posts that it cites and that cite it, and the feeds it was found in. Here's an example for a widely cited post.

      🎓 Expertise Level

      Scour now tries to determine the level of expertise each post assumes and infers the level of expertise you have per topic (based on the wording of your interest is and the types of articles you click on or like). At least for me, this means I'm seeing far fewer beginner Rust questions from Reddit showing up in my feed. (For those in tech, this is powered by Jev.)

      🔎 Search for Feeds and People

      Scour's Search will now show you results for feeds and authors, in addition to posts that match your query.

      Relatedly, you can now follow individual authors as sources and Scour will try to show you their posts from any website they publish on.

      🗑️ Detecting More Junk

      Scour now detects and hides more junk, ranging from sales pages and SEO garbage to uninformative link roundups and low-value discussion threads. You should see more high-quality content in your feeds. By my current count, about 1 in 10 posts being shown before was some kind of junk that Scour now hides.

      ⚡ Faster Feed

      I continue to obsess over making Scour feel super fast and snappy. In September the slowest feed loads got about 7x faster (p99 went from 2.1 seconds to 282 milliseconds) and the median feed load time got 2x faster (p50 went from 90 ms to 40 ms). This is also while ranking about 1.4x as much content as the month before.

      🪦 RIP Reddit Feeds

      Unfortunately Reddit announced their plan to turn off RSS feeds on November 13th. This is how Scour checks which discussions are happening on Reddit and finds articles posted on different subreddits. After November 13th you'll no longer see links to the Reddit discussions from Scour 😢.


      🔖 Some of My Favorite Posts

      Here are some of my favorite articles I found on Scour in September:


      Happy Scouring! - Evan

    3. 🔗 roboflow/supervision supervision-0.30.8 release

      v0.30.8 — Sharper video, cleaner labels

      Video, YOLO labels, masks, VLM parsing and mAP all get more accurate.

      • VideoSink keeps OpenCV's video quality when OpenCV isn't installed.
      • from_yolo reads pose labels as boxes instead of polygons.
      • MeanAveragePrecision scores class-agnostic runs right when only one side has class IDs.
      • from_vlm returns one Florence-2 detection per object, not one per polygon.
      • from_inference masks no longer drift up to a pixel up and left.

      Drop-in upgrade. Without OpenCV, videos get larger; YOLO labels with a negative width or height now raise ValueError.

      ✨ Spotlights / highlights

      sv.VideoSink and sv.process_video keep quality without OpenCV

      The PyAV fallback left the encoder bit rate unset, so mp4v and MJPG files came out at under half of what cv2.VideoWriter writes. It now uses OpenCV's rate settings for every codec except H.264. Files get larger and vp09 may encode more slowly; codec="avc1" keeps files small where an H.264 encoder is available. (#2661)

      import supervision as sv
      
      video_info = sv.VideoInfo.from_video_path("in.mp4")
      with sv.VideoSink("out.mp4", video_info) as sink:  # OpenCV not installed
          for frame in sv.get_video_frames_generator("in.mp4"):
              sink.write_frame(frame)
      # before: mp4v written at under half OpenCV's bit rate, visibly softer
      # now:    same bit rate OpenCV's writer uses
      

      YOLO pose labels load as boxes

      A pose row is a box followed by keypoints. from_yolo used to parse the whole row as a polygon, giving wrong boxes and masks nobody asked for. It now reads the box and skips the keypoints that kpt_shape declares. (#2655)

      ds = sv.DetectionDataset.from_yolo(
          images_directory_path="pose/images",
          annotations_directory_path="pose/labels",
          data_yaml_path="pose/data.yaml",  # kpt_shape: [17, 3]
      )
      # before: polygon-parsed boxes and masks
      # now:    one box per row
      

      Class-agnostic mAP with one-sided class IDs

      A perfect match scored zero when only one side carried class IDs, such as SAM proposals checked against labeled ground truth. With class_agnostic=True, both sides now count as one class.

      One Florence-2 detection per object

      (#2648)

      Florence-2 returns a segmented object as a list of polygons, one per connected region. An object split in two used to come back as two detections; the polygons now merge into one mask with one box around all of them.

      Roboflow masks sit on the right pixels

      (#2649)

      from_inference truncated sub-pixel polygon vertices, shifting each mask up and left by up to a pixel. Vertices are now rounded, the way the COCO, YOLO, LabelMe and Pascal VOC loaders already do.

      🔄 Migration guide

      No migration required for this release.

      📝 Notable changes

      🌱 Changed

      • sv.DetectionDataset.from_yolo raises ValueError naming the annotation file when a label has a negative width or height; it used to load a box with x_min past x_max, which made Detections.area negative and skewed IoU and NMS. as_yolo now orders the corners of a reversed box before measuring, so it no longer writes a file the loader refuses. (#2663)

      🔧 Fixed

      • sv.VideoSink and sv.process_video write mp4v, MJPG and other non-H.264 video at OpenCV's bit rate when OpenCV isn't installed. A frame rate of zero or less now raises RuntimeError in sv.VideoSink, as it does with OpenCV. (#2661)
      • sv.DetectionDataset.from_yolo reads the box of Ultralytics pose labels and skips their keypoints; a kpt_shape other than [K, 2] or [K, 3] raises ValueError. (#2655)
      • sv.DetectionDataset.from_yolo names a malformed annotation line, one with too few values or, for OBB, not nine, in a ValueError instead of failing on an array shape. (#2665)
      • sv.metrics.MeanAveragePrecision(class_agnostic=True) treats detections without class IDs as the same class as labeled ones, and unsigned class ID arrays no longer raise OverflowError on NumPy 2. (#2650)
      • sv.Detections.from_vlm with sv.VLM.FLORENCE_2 merges the polygons of one instance into one detection for <REFERRING_EXPRESSION_SEGMENTATION> and <REGION_TO_SEGMENTATION>, and skips instances with no usable polygon. (#2648)
      • sv.Detections.from_vlm with sv.VLM.QWEN_2_5_VL or sv.VLM.QWEN_3_VL recovers complete detections from a response cut off inside a bbox_2d array or right after a complete object. (#2666)
      • sv.Detections.from_inference rounds polygon vertices to the nearest pixel before rasterising masks, and raises ValueError for NaN or infinite vertices. (#2649)
      • sv.xyxy_to_mask returns an empty mask for a box entirely left of or above the image when its maximum coordinate is a negative fraction. (#2646)
      • sv.Detections.get_anchors_coordinates computes axis-aligned midpoint anchors without integer overflow. (#2660)
      • sv.LineZone.trigger ages crossing history on frames whose detections lack tracker_id, so a reused track ID no longer creates a false crossing after the track expired. (#2644)
      • sv.LineZoneAnnotator(text_orient_to_line=True) no longer raises TypeError without OpenCV for lines drawn right to left. (#2659)
      • The Ultralytics, Inference and YOLO-NAS speed estimation examples measure elapsed time from frame indices; a vehicle missed in one frame of three was reported about 44% too fast. (#2654)
      • examples/speed_estimation/rfdetr_example.py no longer raises AttributeError: supervision._cv2 now provides getPerspectiveTransform and perspectiveTransform. (#2652)

      🏆 Contributors

      • Mohammad Hijjawi (@MohammadHijjawi97, LinkedIn) — fixed video quality without OpenCV, Florence-2 instance merging and Roboflow mask rounding.
      • Kari Pikkarainen (@kari-pikkarainen, LinkedIn) — fixed speed estimation timing, the NumPy flip fallback and the perspective-transform fallbacks.
      • Miral Amin (@aminmiral) — made YOLO loading reject negative extents and name malformed lines.
      • NIKHIL (@Nikhi00718) — fixed LineZone history expiry and anchor overflow.
      • kevin (@kevin9327) — made Qwen parsing recover from cut-off responses.
      • A Aswanth Raj (@aswanth-07, LinkedIn) — fixed class-agnostic mAP.
      • Devulapalli Naga Sri Vaishnavi (@Vaishnavi220506) — fixed masks for off-frame fractional boxes.
      • JANG BYUNGKUN (@8rulerstar) — fixed YOLO pose label loading.

      Automated contributions:@dependabot


      Full changelog : 0.30.7...0.30.8

    4. 🔗 @HexRaysSA@infosec.exchange The upcoming IDA 9.5 adds 3️⃣ new decompilers and will deliver 🔟 platform mastodon

      The upcoming IDA 9.5 adds 3️⃣ new decompilers and will deliver 🔟 platform updates.

      The new decompilers:
      ◾ Android DEX
      ◾ Infineon TriCore
      ◾ Qualcomm Hexagon

      👉 Read the full blog to see the rest of the updates: https://hex- rays.com/blog/ida-9.5-three-new-decompilers

    5. 🔗 Hex-Rays Blog IDA 9.5: 3 new decompilers and 10 platform updates rss

      IDA 9.5: 3 new decompilers and 10 platform updates

      Good tooling starts with solid fundamentals. When every instruction decodes correctly and every function reads as clean pseudocode, you can trust what IDA shows and put your time into the binary itself. That holds whether the one reading the output is a person or an agent driving IDA. IDA 9.5 brings that reliability to three new architectures and sharpens it on several familiar ones.

    6. 🔗 Andrew Ayer - Blog sourcespotter-authorize: Monitor Your Go Modules for Malicious Versions, Without the Noise rss

      You can protect the users of your Go modules from supply chain attacks, such as a compromise of your GitHub account, by monitoring Go's checksum database (sumdb). Since the go command won't install a module unless its checksum is published in the sumdb, monitoring the sumdb lets you discover unauthorized versions of your modules. Since the sumdb is a transparency log, you can even detect if Google themselves go rogue and publish a malicious version of your module. Although you can only detect, not prevent, attacks, Go's Minimal Version Selection makes it possible to respond before most of your users have installed the malicious version. Dependency cooldowns, which are coming to Go, will make it even easier to respond in time.

      How To Monitor

      Source Spotter, which is operated by my company SSLMate as a free service to the Go community, provides Atom feeds listing all versions of your modules found in the sumdb. For example, this Atom feed returns all versions of modules under the src.agwa.name/ prefix:

      https://feeds.api.sourcespotter.com/modules/versions.atom?module=src.agwa.name%2F

      You can also get Prometheus-compatible metrics:

      https://metrics.api.sourcespotter.com/modules?module=src.agwa.name%2F

      You can scrape the metrics endpoint using Prometheus and alert in your usual way, subscribe to the Atom feed using your favorite feed reader, or use one of the free services that converts Atom feeds to emails.

      Avoiding Noise

      Discovery is only half the story. The more important half is deciding whether to alert on a discovered module. Approximately 100% of the records published in the sumdb are legitimate. If you're alerted every time a new version of your module is legitimately published, it will be very hard to notice the one time it's an attack.

      I wasn't sure at first how Source Spotter could facilitate no-noise monitoring. Initially, I thought Source Spotter would need access to your module's Git repository so it could cross-check sumdb records against the repo's contents. But that seemed complicated, and it would fail to detect a compromise of your repository host. I also really wanted a solution that wouldn't require users to create accounts.

      I finally found a lightweight solution that I really like. I'll show you how you use it, and then explain how it works.

      First, you install a small command line tool called sourcespotter- authorize and generate a public/private key pair:

      $ go install software.sslmate.com/src/sourcespotter/cmd/sourcespotter- authorize@latest sourcespotter-authorize -keygen

      Run sourcespotter-authorize again with the -feed-for or -metrics-for flags to output Atom and Prometheus URLs for the module prefix you want to monitor (src.agwa.name/ in this example):

      $ sourcespotter-authorize -feed-for src.agwa.name/ https://feeds.api.sourcespotter.com/modules/versions.atom?module=src.agwa.name%2F&mldsa=efbcb2bcb2d4decdf1ad9cab3224b2dd0087b6c6877abe00448a00d31716dff1 $ sourcespotter-authorize -metrics-for src.agwa.name/ https://metrics.api.sourcespotter.com/modules?module=src.agwa.name%2F&mldsa=efbcb2bcb2d4decdf1ad9cab3224b2dd0087b6c6877abe00448a00d31716dff1

      These are the same URLs shown earlier, but with a new mldsa= parameter in the query string which is the hash of the public key you generated in the previous step.

      Initially, these URLs return the same contents as the URLs without the mldsa= parameter. That's because you haven't marked any module versions as authorized yet.

      To mark a module version as authorized, change into the Git repository for the module and run sourcespotter-authorize with a Git tag:

      $ cd ~/src/snid $ sourcespotter-authorize v0.4.0

      Now, v0.4.0 of this module is omitted from the feed and metrics URLs.

      The command accepts multiple tags as arguments, so you can authorize all the tags in a repo like this:

      $ sourcespotter-authorize $(git tag)

      Moving forward, you should run sourcespotter-authorize any time you tag a new version. I've written a tiny shell script called gotag that runs git tag followed by sourcespotter-authorize:

      #!/bin/sh -e git tag "$1" sourcespotter-authorize "$1"

      As long as you authorize every tag you create, the feed and metrics URLs will report zero unauthorized module versions, eliminating false positive alerts. A supply chain attacker can't hide their malicious versions from your feeds unless they also compromise sourcespotter-authorize's private key.

      How It Works

      sourcespotter-authorize -keygen generates an ML-DSA-44 private key (stored under $XDG_CONFIG_HOME/sourcespotter-authorize). The SHA-256 hash of the corresponding public key goes in the mldsa query string parameter.

      When you authorize a tag, sourcespotter-authorize uses the golang.org/x/mod/zip and golang.org/x/mod/sumdb/dirhash packages to compute the checksums of the go.mod and module zip files for the tag. It formats the module path, version, and hashes for each tag as a go.sum file (no need to invent a new format here), and signs the file with your private ML-DSA key. Finally, it uploads a JSON object to the Source Spotter server containing your public key, the go.sum file, and the signature.

      The Source Spotter server verifies the signature using the public key. If valid, it records the module path, version, and hash as being authorized by that public key, and excludes that module version from the feeds and metrics endpoints for the key.

      The protocol is simple and documented if you want to implement your own client.

      Why Not Just Sign Your Releases?

      If you're generating a private key and signing stuff anyway, why not just sign your releases and distribute the signatures? After all, this is the traditional approach to stopping unauthorized software releases.

      The problem with the traditional approach is that consumers of your software have to verify the signatures, which means they need to know what your public key is. Key distribution is a hard, hard problem, and most software ecosystems do a bad job at it. I would guess that in practice, few people actually verify signatures of software releases.

      In contrast, the go command automatically verifies that every module it installs is listed in the checksum database, without the user needing to do anything. Although sourcespotter-authorize needs a private key, you only need to distribute the public key to Source Spotter (or whatever other monitor you might choose to use), rather than every single consumer of your software. That's a vastly simpler problem, especially when (not if) you need to rotate your keys.

      Everyone who signs their software releases will need to rotate their keys soon, as quantum computers are expected to break RSA and elliptic curves within the next few years. An earlier version of sourcespotter-authorize used elliptic curve keys. After it gained ML-DSA support, upgrading my key was super easy - generate a new key, transfer my list of authorized versions to it, and finally update the URL in my feed reader:

      $ sourcespotter-authorize -export > tmpfile $ rm ~/.config/sourcespotter-authorize/private_key $ sourcespotter-authorize -keygen $ sourcespotter-authorize -import < tmpfile $ sourcespotter-authorize -feed-for src.agwa.name/

      This was so much easier than communicating a new key to every consumer of my software would have been, and didn't require a lengthy transition period during which I was signing with two keys.

      You can learn more about Source Spotter's module monitoring, get the source code on GitHub, or read my blog post about how Source Spotter also verifies Go's reproducible builds.

    7. 🔗 backnotprop/plannotator v0.28.5 release

      Follow @plannotator on X for updates

      Missed recent releases? Release | Highlights
      ---|---
      v0.28.4 | Comments come back after the agent edits a file, the agent can list and close its reviews, PR-description feedback no longer dropped, Pi thinking levels
      v0.28.3 | Typing into your agent while it answers an Ask no longer streams into Plannotator
      v0.28.2 | Question cards show wrapped choices and tables, pinned images named by their file, Done with nothing to send starts no agent turn
      v0.28.1 | Ask AI from diagram comments, pinned images named for the agent, OpenCode 1 URL toasts and one reply per feedback, folder feedback sent once
      v0.28.0 | Ask this session in Claude Code, Pi and OpenCode 2, Claude Code mod on by default, Pi plan review no longer blocks, first-run demo
      v0.27.25 | Code review works with color.diff = always, Bitbucket review fixes, wide tables no longer collapse in Firefox, install script fix
      v0.27.24 | Image previews stay in the all-files view, PR comment previews open on the commented line
      v0.27.23 | Bitbucket Cloud PR review, Question UI for answering agents in place, opt-in auto-update, viewed files remembered, review another repo or worktree
      v0.27.22 | Plans open the docs they link to, Claude review jobs locked down, Code Tour on Linux without Claude's sandbox, Pi reviews the latest plan
      v0.27.21 | Remote and phone sessions load several times faster, real Request changes on GitHub, model pickers show real names, OpenCode fixes
      v0.27.20 | Mistral Vibe support, annotate gets the full Options menu and Settings, jj Commits panel, long lines wrap in plan code blocks

      What's New in v0.28.5

      Twelve PRs, one from a first-time contributor. Agents on Pi and OpenCode 2 can now open Plannotator reviews themselves without holding the session, a review can cover several files at once, and every decision now names the exact file it is about.

      Several files in one review

      plannotator annotate spec.md ui/mock.html notes.md now opens one review of all the files, in the order you typed them. A header switcher shows "2 of 3", the Files tab keeps that order, and one decision covers the whole set, with the feedback split into a section per file. Ask AI knows which file you are on, and your unsent comments are kept per file. The same works when an agent passes a list of files to the plannotator tool. If one of the paths does not exist, nothing opens and the error names the missing file. Words that are not file paths keep their old meaning, so annotate look at notes.md please still opens notes.md. (#1718)

      The plannotator tool on Pi and OpenCode 2, off until you turn it on

      On Claude Code (with the Plannotator mod) agents already had a plannotator tool. It is now on Pi and OpenCode 2 too. It matters when the agent opens Plannotator itself, for example when you ask it to "show me the HTML plan in Plannotator" or "open a Plannotator code review". Without the tool, the agent runs the plannotator command in its shell, which blocks the session until you finish and leaves Ask AI unable to reach the session. With the tool, the review opens right away while the agent keeps working, your decision comes back as a message, Ask this session works from that review, and the agent can list and close the reviews it opened.

      The tool adds about 780 tokens to every request, so on Pi and OpenCode 2 it is off by default. The first Plannotator page you open there asks "Do you use Plannotator as a skill?" and turns it on for your next session if you say yes. You can change it any time with the new toggle in Settings → General (plan review, annotate and code review), the agentTool key in ~/.plannotator/config.json, or PLANNOTATOR_AGENT_TOOL. On Claude Code the tool stays on by default, since Claude loads it only when it is needed; the same switch turns it off there without turning off the rest of the mod. Your /plannotator-* commands work the same either way. (#1714, #1715, #1724, #1725)

      The tool's description no longer includes the unfinished reply action, and Pi's bundled knowledge skill stays out of the model's context by default, keeping the promise from #842.

      Every decision names the file it is about

      A user reviewing two different files that were both named QUESTIONS.md approved one of them. The approval reached the agent as just "QUESTIONS.md — Approved." with no path, and the agent attached it to the other file. Every decision message now carries a Target: line with the full path (or URL, PR, or folder), taken from the review that recorded the decision, and reviews that share a file name are labelled with their folder, such as releases-2026-10-04/QUESTIONS.md. A code review decision names the PR that is on screen when you decide, even if you switched PRs inside the review.

      A related gap is closed too: when a review's port was reused (a fixed PLANNOTATOR_PORT, remote mode, or rarely by chance), an old browser tab could submit a decision to the new review. Each review now has its own id, and a decision from a tab that belongs to a different review is refused with a "This review was replaced, reload" banner. plannotator sessions now lists each review's id and full path, and has a --json option. (#1729)

      Ask this session reconnects after sleep

      With the Claude Code mod, Ask AI in an open review could switch permanently to "This session is no longer available" after your computer slept or the network dropped for a few minutes. It now reconnects on its own, waiting a little longer between attempts while the review is unreachable. Running claude --continue while the old window is still open no longer makes the two windows fight over the review: the window you are using takes it over, each decision is delivered exactly once, and plan approvals reach the right window. If you quit Claude Code after a decision arrived but before Claude saw it, you get a notice next time saying where it was saved. (#1727)

      Approve with a note in every gated review

      Gated annotate reviews (--gate, or the tool with gate: true) opened from Claude Code never offered "Approve with a note…", even though the note could be delivered. They do now, for markdown and HTML alike, on every host that delivers the note with the approval. Your note reaches the agent as an "approved with notes" message. A plain approval is unchanged: plain output is still exactly The user approved. One thing for scripts: when you approve with a note, plain output now prints the approved-with-notes message instead of that line. Scripts should use --gate --json. (#1728)

      Additional Changes

      • Done says nothing was sent. Clicking Done with nothing to send showed "Feedback Sent". It now shows "Done, nothing was sent", and Close no longer claims a response was sent (#1730).
      • Pi keeps the model you picked. Using /tree to re-answer something during plan execution switched the model back to the plan's original one. It now keeps your choice; moving into or out of plan mode still applies the phase's model (#1723, fixes #1722).
      • Gruvbox diffs. Added lines in code review were grey on the Gruvbox theme. They are now Gruvbox green in light and dark (#1721).
      • Safer settings. A web page in your browser can no longer change Plannotator's settings in the background, and saving settings from the VS Code panel works again, as do viewed-file progress and the Call Flow install there (#1724).
      • Type checking now covers the Claude Code side's server code (#1720).

      Install / Update

      macOS / Linux:

      curl -fsSL https://plannotator.ai/install.sh | bash
      

      Windows:

      irm https://plannotator.ai/install.ps1 | iex
      

      Claude Code Plugin: The plugin and the plannotator binary update separately, so run the install script above as well. In a terminal:

      claude plugin marketplace update plannotator
      claude plugin update plannotator@plannotator
      

      Then restart Claude Code. Inside Claude Code, run /plugin marketplace update plannotator, then open /plugin → Installed → plannotator → Update now.

      Pi:

      pi update --extensions
      

      OpenCode: Re-run the install script above. It now also clears the OpenCode 2 plugin cache.

      What's Changed

      • feat(annotate): several files in one annotate review by @backnotprop in #1718
      • feat(pi): the plannotator tool on Pi by @backnotprop in #1714
      • feat(opencode): the plannotator tool on OpenCode 2 by @backnotprop in #1715
      • feat(tool): agentTool switch with per-host defaults, Pi tool stability, drop the reserved reply action by @backnotprop in #1724
      • feat(ui): agent tool switch in Settings and a one-time offer on Pi and OpenCode 2 by @backnotprop in #1725
      • fix: decisions name their full target; refuse stale-tab decisions on a reused port by @backnotprop in #1729
      • fix(mod): restart a dead Ask-this-session bridge; one watcher per launch across processes by @backnotprop in #1727
      • fix(annotate): offer Approve with a note in every gated session that delivers it by @backnotprop in #1728
      • fix(annotate): show a Done screen, not Feedback Sent, when Done sends nothing by @backnotprop in #1730
      • fix(pi): keep the user's model on a /tree navigation within the same phase by @backnotprop in #1723
      • Fix gruvbox positive diff colors by @TheEdgeOfRage in #1721
      • fix(hook): type-check apps/hook/server and fix vibe-plan.ts narrowing by @backnotprop in #1720

      New Contributors

      Contributors

      @TheEdgeOfRage fixed the grey added lines in Gruvbox code review, with the color override the colorblind theme already uses.

      Community:

      Full Changelog : v0.28.4...v0.28.5

    8. 🔗 Project Zero How to fix a bug in a fix rss

      Project Zero often works with software vendors to remediate the vulnerabilities we report and provide broader guidance on making software more secure. Some vendors express concern about potential scenarios in which they are unable to fix vulnerabilities that are causing immediate user harm, due to limitations in their patch delivery systems. Since Project Zero encounters a wide array of systems designed to protect users in the case of exceptional exploitation scenarios, both through vendor discussions and security reviews, we want to share what we’ve learned.

      This post provides an overview of systems in use by large vendors that allow them to remediate small volumes of vulnerabilities much faster than their typical update process. Our goal is to provide a reference for vendors seeking to implement or enhance the capabilities of such systems, and to encourage vendors to consider how they would fix an urgent vulnerability before they receive one.

      Why patching takes time

      Patching a vulnerability typically involves the following stages:

      • Triage — a vulnerability report is received, validated, prioritized and assigned to a specific developer to be fixed
      • Patch development — a software development team writes, reviews and commits code that fixes the vulnerability
      • Testing — the patch is tested to ensure the vulnerability is remediated and the software still functions correctly when the patch is applied. This can include formal testing by a test team, automated testing and alpha and beta testing where a patch is shipped to a limited group of users for feedback on normal use.
      • Partner review — some software updates require review by third parties before they can be shipped, due to relationships between the software vendor and other organizations, for example, carrier acceptance for some mobile updates.
      • Delivery — the patch is delivered to and installed by end users
      • Activation — sometimes an additional step, such as a system restart, is needed to switch the system to the updated software

      Of course, this is a simplified picture. Patching can involve repeating steps, for example rewriting a patch if tests fail, or additional stages when third- party vendors are involved. However, this is a minimal set of steps most software updates require.

      The challenges of emergency patches

      While triage and patch development time contribute substantially to the speed at which vendors can generally patch vulnerabilities, they contribute less to emergency patch time. Triage is usually very fast in situations where vendors know they have an urgent problem, and patch development can be expedited based on priority. Only in rare circumstances, where a vulnerability is especially complex, or a vendor’s security team does not have a complete picture of their software’s components and who within their organization maintains them, have we seen urgent patches delayed in the triage or development phase. Likewise, partner agreements usually have exceptions for updates in emergency situations.

      Most vendors’ patch speed is limited by the testing and delivery stages. Testing is important because all changes to software risk introducing unexpected behavior. The worst-case scenario is that inadequately tested software ‘bricks’ a device, causing it to malfunction in a way that it can no longer perform key functionality or receive software updates to remediate this. Buggy software updates have also led to situations where user data is corrupted or lost, and any decrease in software functionality after a security update makes users less likely to apply updates in the future.

      The potential cost to vendors of shipping poorly tested updates varies depending on the nature of the underlying software. For example, if a mobile application is rendered unusable due to an update that corrupts local data or prevents it from launching, users can easily install the next version via an app store, and their data is usually saved on a remote server, so costs are limited to user support. Meanwhile, if a mobile device gets bricked, it needs to be returned to its manufacturer or place of purchase for repair, leading to substantial costs for the vendor and potentially the user.

      The possibility of serious functional bugs is considered in the design of most patch delivery systems. Updates are often rolled out slowly, so that serious problems can be detected before they affect too many users. Often, patching vulnerabilities quickly and avoiding buggy patches are at odds with each other, requiring tradeoffs that prioritize one over the other.

      A variety of other technical challenges can limit the speed of patch delivery. One is the design of the patching system. A common design is that devices probe for updates at a regular interval, leading to patch saturation being limited to that interval. ‘Push’ style update systems can deliver patches to all users faster, but generally require more infrastructure.

      User behavior and environment can also be a barrier to patch propagation. Patches that require user interaction to install are often delayed by users, and network speed and data cost are also factors in installation rate. Updating many users at once, as opposed to over a period of time, can strain patch delivery infrastructure. Chrome and Microsoft have written about the challenges of updates requiring restart to install, as users are often reluctant to restart their system and restarts take time.

      While testing delays and limitations of the patch delivery system affect all updates, the shorter time frame of emergency updates make them a larger contributor to the overall time it takes to deliver a patch.

      Emergency patching methods

      Feature flags

      Feature flags are conditional statements in source with paths determined by values provided by a remote server. They are often used for A/B testing, but they can also be used for short term remediation of vulnerabilities in emergency situations. A widely publicized case of this was a serious 2019 FaceTime vulnerability, where Apple temporarily disabled Group Facetime with a feature flag. Several vendors have made at least some media codecs available in 0-click contexts controllable via feature flags, and can disable them in the case of active exploitation, falling back to another codec for realtime transmission.

      The main benefit of feature flags as a vulnerability remediation method is that testing can be performed with each flag set in advance, so a fast update does not require shipping untested code. They can also be delivered to users much more quickly, as updating feature flags requires transmitting a very small amount of data.

      Recently, Meta published a blog post on how they implemented a ‘dual stack’ library, in which two versions of the WebRTC video conferencing library were compiled into a single binary, with the version in use controllable via a feature flag. This technology enables rapid updates with less testing, as new versions can be shipped with the option to quickly move users back to the previous version if function problems occur. While Meta uses two versions of the same library, it would also be possible to create a ‘dual stack’ with two different libraries that implement the same features (for example, two H264 libraries), allowing an application to switch to a different library to render a specific vulnerability unreachable without loss of functionality in an emergency. This would require additional testing, but it is testing that can be performed up front. It could also be possible to have a second library that enables performance intensive mitigations that would block many possible bugs, such as ASAN, or enabling DCHECKs.

      Filtering

      Filtering is running a dynamically updatable ruleset, such as a regular expression, against untrusted input in order to block specific input that is required to reach a vulnerability. An example of this is Android’s Intent Firewall, which allows specific usages of an Android IPC mechanism called intents to be disabled based on rules in a dynamically updateable XML file, which enables blocking intents that can be used to exercise specific vulnerabilities. It was recently used to block vulnerabilities in third-party Android wallets.

      Some platforms have endpoint detection software that can perform filtering on a wide variety of system input, for example Microsoft Defender on Windows systems, and Google Play Protect on Android devices. Rules that block specific exploits or make certain vulnerabilities unreachable can often be deployed to these applications very quickly. Endpoint detection requires parsing a great deal of untrusted input, often in privileged context, so these applications are not without risk, but in systems where they already exist, they are a potential method of emergency remediation.

      As an approach, filtering is more flexible than feature flags. For feature flags to be effective, the vendor needs to determine what features they might want to disable in advance, and if this isn’t comprehensive, they might find themselves in a situation where a vulnerability can’t be remediated via feature flags. Meanwhile, filtering can be used to block a wide variety of inputs, even ones that have never been considered. The downside of filtering is that performing filtering frequently can decrease software performance, and at least some testing of new filters is required, and can’t be performed upfront without knowing the vulnerability that needs to be blocked, as it is possible to write filters that interfere with necessary system functions.

      Alternate Channels

      The network ‘channels’ used to deliver software updates to users can be slow for a variety of reasons discussed above. Vendors sometimes implement alternate channels that can be used to deliver smaller updates more quickly.

      Android Pony Express (APEX) is an example of an alternate channel that can be used to ship updates to specific high-risk Android components faster than a full system update. It shortens the patch development time, as OEMs do not need to integrate updates to APEX components. APEX is available to OEMs, and can be used to update OEM- maintained libraries.

      Several applications we’ve researched have the ability to update individual libraries outside regular updates, usually by having some flag that is regularly checked over the network, and then downloading the library and loading it with dlopen or equivalent. While this is an effective way to avoid delivery-speed limitations of updates, it can also introduce critical vulnerabilities if libraries delivered in this way are not adequately verified by the client to have originated from the vendor. We encourage vendors to be cautious, and ensure that emergency update mechanisms of this variety have adequate security testing.

      Hotpatching

      Some vendors have implemented update mechanisms that allow units of binary code smaller than libraries to be delivered and applied directly to the memory space of a running process. For example Linux supports Livepatch which enables kernel functions to be directly replaced in memory without a restart. Similarly, Windows’ hotpatch allows security updates that contain only updated functions to be delivered to users, and applied while the process is still running.

      Hotpatching has the potential to deliver very flexible security patches to software very quickly, with no degradation of user experience, though it typically has some limits to the nature of patches it can deliver, for example, updates that require changing the definition of a structure shared between functions are sometimes not supported. Hotpatching has similar security downsides to alternate channels, and also carries the risk of introducing ways to bypass exploit mitigations, as it requires permissions to map pages with write-execute privileges at some point during patching. It also doesn’t address any of the testing challenges of rapid updates, just the delivery challenges.

      The importance of emergency patching

      LLMs are increasing the vulnerability discovery and exploitation capabilities of both attackers and defenders. A wider array of actors now have the ability to perform novel attacks at greater speed. In light of this, it is important for vendors to consider how to protect their users in the case of active exploitation. Rapid update mechanisms do not need to be heavyweight or be capable of fixing every possible bug and preserving perfect user experience in every scenario. Technologies like feature flags, filtering and alternate update mechanisms can remediate the most likely and severe vulnerabilities in the short term, while keeping devices reasonably functional for users.

      It is urgent for vendors to plan how they will protect their users in the worst case scenario of widespread active exploitation. Actions taken now can greatly improve security outcomes for users. By taking stock of update mechanisms already available to them and implementing rapid remediation functionality where gaps exist, vendors can be better prepared for whatever the future holds.

    9. 🔗 syncthing/syncthing v2.1.6 release

      Major changes in 2.1

      • Devices and folders can now be grouped in the GUI by setting the new
        group attribute.

      • HTTP and HTTPS proxies with support for CONNECT can now be used, in
        addition to the existing support for SOCKS proxies (the environment
        variable all_proxy=https://...).

      • Block indexing can be turned off for folders where it's more desirable to
        optimise for reduced database size and overhead than minimal transfer
        size (the blockIndexing attribute on folder configuration).

      • GUI login session duration can be configured to be longer or shorter than
        the default one week, or set to infinitely long. The cookie path can also
        be adjusted. (The sessionCookieDurationS and sessionCookiePath
        attributes in the GUI configuration.)

      This release is also available as:

      • APT repository: https://apt.syncthing.net/

      • Docker image: docker.io/syncthing/syncthing:2.1.6 or ghcr.io/syncthing/syncthing:2.1.6
        ({docker,ghcr}.io/syncthing/syncthing:2 to follow just the major version)

      What's Changed

      Fixes

      • fix(model): introducers should not be able to add themselves to folders by @calmh in #10880
      • fix(gui): add accessible labels to buttons in Edit Device modal (fixes #10873) by @tomasz1986 in #10874
      • fix(model): properly error out when a temp file can't be created by @calmh in #10883
      • fix: disable keepalive on most outgoing HTTP connections by @calmh in #10891
      • fix(monitor): continue log writes when stdout is unavailable on detached Windows consoles (fixes #10882) by @Shablone in #10889
      • fix(gui): improve header and button color contrast (fixes #10488) by @giri256 in #10815
      • fix: use global short lived HTTP client with proper HTTP/2 support by @calmh in #10901
      • fix: remove unnecessary delay in index transmission by @calmh in #10908
      • fix(watchaggregator): properly convert floating-point time duration (fixes #10899) by @calmh in #10900

      New Contributors

      Full Changelog : v2.1.5...v2.1.6

    10. 🔗 streamyfin/streamyfin v0.55.1 release

      fix(ios): keep tab labels on their tabs when built with the iOS 27 SD…

    11. 🔗 HexRaysSA/plugin-repository commits sync repo: +1 release rss
      sync repo: +1 release
      
      ## New releases
      - [ida-settings-editor](https://github.com/williballenthin/ida-settings): 1.3.1
      
    12. 🔗 Mitchell Hashimoto A Terminal Protocol for Program Status (OSC 7501) rss
      (empty)
    13. 🔗 Filip Filmar Razboj: a minimal GPU in TxHDL rss

      Razboj is a minimal graphics rasteriser implemented in approximately one hundred lines of TxHDL. It reads a display list from memory and writes rendered pixels into a framebuffer over an AXI bus. TxHDL lowers the design to synthesizable Verilog and VHDL, and the build verifies both netlists against the software simulation trace. This post describes the rasteriser architecture and hardware design tradeoffs.

      What it draws

      The reference demonstration scene measures 64 by 64 pixels (4,096 pixels total). The scene consists of eight display list entries: a background clear, three rectangles, and four triangles. The rasteriser renders the entire scene in approximately 13,000 clock cycles. A verification harness reads the completed framebuffer from memory and saves the output image.

    14. 🔗 Armin Ronacher What is Codemode rss

      More than a year ago I wrote a few posts here that recommended people not to load custom tools into their context (or MCP servers) but to just use more scripts. Most importantly I wrote that Code Is All You Need and I wrote about that MCP needs code. With Pi 1.0 we now added MCP support via Codemode which in some ways is a long time coming, but then also maybe somewhat surprising to some. So I want to share some updated thoughts on this blog on what this all means.

      What Are Tools

      When a harness like Pi provides tools for an LLM to call, it does so by supplying some tool definitions which then translate into some token structure on the server side. Whether a model is encouraged to call a tool is the result of the reinforcement learning process. Something I wrote about before if you want to learn more.

      One of the reasons we strongly lean towards CLI and bash is because it allows easy composition of calls, and because the model also learns how the file system works when it's trained. So when it invokes a tool like echo foo > /tmp/test.txt the model also learns that after that tool call, there is now a file called test.txt in /tmp.

      However bash has one fundamental limitation which is that it can only compose programs that run. And there are some things, which are not programs, but native tools to the LLM and they sort of have to be.

      The most obvious example here is read or view_image. If a multimodal model needs to read an image, it cannot use cat for that because the harness needs to inject the actual image payload into the protocol of the LLM.

      Another quite vivid example are sub agents. In order to spawn and orchestrate sub agents, it's tricky to avoid tools that are provided by the harness. While in theory the agent could provide a CLI tool that talks to the outer harness via environment variables and Unix sockets, it's a rather crude process. It however has another issue, and that is where the code runs.

      Brains vs Hands

      To better understand that, it's important to think a bit more about where all the bits and pieces run. There really usually are two different systems involved. The first is the brain, the harness: it runs on one machine. It's trusted. The second is often the same machine, but it's really where the tools are executing: the hands. In Pi we now call this the execution environment, but you can think of it as the target of all the operations.

      Crucially what is important for us, is that there is a dividing line between the harness brain and the target environment that runs bash and executes the tools.

      And splitting this in half has some really important consequences. For a start it means that they are running on different file systems and they have different levels of trust. If you for instance use a sandboxing solution like Gondolin your bash stuff will be sandboxed just fine, but the harness itself will not be.

      Orchestrating The Harness

      Which brings us to what Codemode really does: it's a way for the LLM to express and orchestrate complex operations on the harness side, but not the execution environment side. Codemode runs in the harness, in its own sandbox. In case of Pi it's running in QuickJS within a WASM runtime with intentional limitations: no network, no file system, no timers, limited RAM. The only way is to call more tools. You could also imagine that Codemode could run Scheme or some other language as well.

      If you are not familiar with Codemode, it's basically just a way to issue tool calls from within some language, in our case JavaScript. That allows you to compose those calls without necessarily going through the LLM's context. Credit for naming goes to our friends at Cloudflare who coined it.

      For instance if you issue a bash call as a regular tool call in the LLM, then we only throw the trailing 2000 lines into the context and if the agent wants more, it needs to look at the overflow file itself. If however the agent issues that invocation via Codemode, then the Codemode side gets larger outputs sent structurally.

      Most importantly, because Codemode is JavaScript the agent can express concurrent operations and basic workflows. A common way in which you see agents now use this, is to first probe at 5-10 items from some tool response to see what it looks like, and to then write a Codemode script that processes the next n items.

      Codemode also allows you to throw state into the transcript! That means that one Codemode invocation can stash away data, that the next call in the session can load again. And remember: this is on the harness host, not the sandbox.

      In case of Pi, Codemode also allows you to issue calls that naturally do not make any sense in Pi's traditional interface. For instance if you want to generate images with an image model or you want to classify some text with a one shot classifier model, those Pi APIs are exposed via Codemode, but not via regular tools where they would just waste context.

      What It Looks Like

      So now that we talked a bunch about it, it's probably worth being a bit more explicit about it. Let's walk ourselves through some invocations of Codemode of recent Pi sessions of mine. Note that none of this code is human written. It's from real sessions of Pi, just re-indented for your viewing pleasure. The agent starts using Codemode automatically either because it's a task where the model already naturally picks up that tool, or because a user asked it to.

      Note that Codemode is by default only enabled in Pi when MCP is enabled, but you can turn it on with "defaultTools": ["+codemode"] in the settings. Just ask Pi to enable it for you.

      Generating Images

      Let's start simple with image generation. Image generation is a feature that Pi supports in the AI SDK core, but it's not a tool that the agent can use. In the past the only way to use image models has been to write a bespoke extension or to have the agent run node itself and use the internal image APIs. However because we expose quite a few of the internal model APIs within Codemode, it means that the agent can use it:

      const [painter] = await models.getAvailableOfType("image");
      const result = await models.generateImages(painter, {
        input: [{ type: "text", text: "A cute little puppy sitting on a grassy " +
          "lawn, soft natural light, photorealistic" }],
      });
      if (result.stopReason !== "stop") return result.errorMessage;
      
      for (const block of result.output) {
        if (block.type === "image") image(block);
        else text(block.text);
      }
      

      Note that the call to image() sends the image back as image content to the LLM. On the harness side it feeds it directly into both the agent, as well as onto disk as a temporary artifact in case the agent wants to be able to pass that image back to bash.

      Classifying Things

      Similar things apply to classifier models such as Jev. They also do not fit well into the workflows of an agent through the typical tools. But rather than making a bespoke tool available, Codemode just allows the agent to reach into the AI SDK and invoke those directly. Here you can see how Jev is used to mass process GitHub issues for a quick sentiment analysis:

      const jev = await models.getModelOfType("classifier", "typesafe", "jev-latest");
      const r = await tools.bash({
        command: "gh issue list --state open --limit 100 " +
          "--json number,title,body,comments",
      });
      const issues = JSON.parse(r.output);
      
      const results = await Promise.all(issues.map(async (issue) => {
        const res = await models.classify(jev, {
          state: {
            title: issue.title,
            body: (issue.body || "").slice(0, 4000),
            comments: issue.comments.slice(-5).map(c => c.body.slice(0, 800)),
          },
          questions: {
            sentiment: {
              type: "choice",
              instructions: "What is the overall sentiment of the author towards pi?",
              criteria: {
                positive: "Appreciative, happy, constructive praise",
                neutral: "Matter-of-fact report or request without emotion",
                negative: "Frustrated, annoyed, upset, or angry",
              },
            },
            frustration: {
              type: "score",
              instructions: "How frustrated is the reporter?",
              criteria: ["not at all", "mildly", "clearly frustrated", "very angry"],
            },
            kind: {
              type: "choice",
              instructions: "What kind of issue is this?",
              criteria: {
                bug: "Bug report or regression",
                feature: "Feature request or enhancement",
                question: "Question or support request",
                other: "Docs, discussion, meta, spam",
              },
            },
          },
        });
        if (res.stopReason !== "stop") {
          return { n: issue.number, title: issue.title, error: res.errorMessage };
        }
        return { n: issue.number, title: issue.title, ...res.answers };
      }));
      
      store("sentiment_results", results);
      return results
        .filter(r => !r.error)
        .sort((a, b) => b.frustration.score - a.frustration.score)
        .slice(0, 12)
        .map(r => `#${r.n} ${r.frustration.score.toFixed(2)} [${r.kind.choice}] ${r.title}`);
      

      Note how in that above example we also call store() which dumps the result of that execution into the session transcript. A future invocation of Codemode can thus read back that result if it wants to.

      The Promise.all here is fine, because Pi limits the total number of concurrent tool executions itself to four and maintains a queue for the rest.

      A more adventurous example is to use Jev to drive a game engine for debugging purposes:

      Codemode with Jev for Game Debugging

      Here it knows about my tankctl command and it built itself quickly a minimal harness around it to drive a game loop to assist a user with debugging a problem. Note how it built a 30 step loop in which each step goes back to both the game engine to get a text dump of what's going on, and then to Jev to determine what to do next:

      const jev = await models.getModelOfType("classifier", "typesafe", "jev-latest");
      const tank = async (cmd) =>
        (await tools.bash({ command: `tools/tankctl "${cmd}"` })).output;
      await tank("start --map assets/maps/night_arena.map");
      
      const questions = {
        action: {
          type: "choice",
          instructions: "You control the tank '@' in a top-down tank game. " +
            "Choose the best next action.",
          criteria: {
            attack: "an enemy has line of sight to you and you can fire at it",
            approach: "no enemy has line of sight; drive toward the nearest enemy",
            dodge: "an enemy shot is heading at you and will hit soon",
            powerup: "a powerup is close and no enemy threatens you",
          },
        },
      };
      
      function commandFor(choice, st) {
        const p = st.player;
        const enemy = st.enemies.filter(e => !e.dead)
          .sort((a, b) => (b.los - a.los) || (a.dist - b.dist))[0];
        if (choice === "attack" && enemy) {
          return `fire_at tank ${enemy.id}; frames 30 until clear,damage,kill`;
        }
        if (choice === "dodge") {
          // move perpendicular to the closest incoming shot
          const s = st.projectiles.filter(s => !s.yours)
            .sort((a, b) => a.eta - b.eta)[0];
          const dir = s && Math.abs(s.vel[0]) > Math.abs(s.vel[1])
            ? (p.pos[1] > s.pos[1] ? "+down" : "+up")
            : (p.pos[0] > (s ? s.pos[0] : 0) ? "+right" : "+left");
          return `input ${dir}; frames 20 until damage; input stop`;
        }
        const powerup = st.powerups.filter(u => u.available)
          .sort((a, b) => a.dist - b.dist)[0];
        if (choice === "powerup" && powerup) {
          return `goto ${powerup.pos[0]} ${powerup.pos[1]} 180`;
        }
        return enemy ? `goto ${enemy.pos[0]} ${enemy.pos[1]} 90` : null;
      }
      
      const log = [];
      for (let step = 0; step < 30; step++) {
        const st = JSON.parse(await tank("state"));
        if (st.state !== "playing") break;
        const threats = st.projectiles
          .filter(s => !s.yours && s.miss_dist < 1.5 && s.eta < 1.5)
          .map(s => `incoming shot dist ${s.dist} eta ${s.eta}s`)
          .join("\n") || "no incoming shots";
        const r = await models.classify(jev, {
          state: { map: await tank("view 8"), threats, hp: st.player.hp },
          questions,
        });
        if (r.stopReason !== "stop") {
          log.push(`#${step} classifier error: ${r.errorMessage}`);
          break;
        }
        const choice = r.answers.action.choice;
        const cmd = commandFor(choice, st);
        if (!cmd) break;
        log.push(`#${step} hp=${st.player.hp} ${choice} -> ${await tank(cmd)}`);
      }
      return log.join("\n");
      

      Calling MCP Servers

      Lastly, Codemode obviously is great for calling MCP servers. And because we do not actually expose any of the MCP tools to the LLM, the agent first uses provided APIs to issue a tool search within Codemode to discover what it might be able to do with the connected servers. This form of progressive discovery makes the whole MCP business work well enough for a lot of use cases today.

      Here for instance you can see the agent reach for the Sentry MCP straight away, even without discovering the tools, presumably because it has learned during the RL process already about what the Sentry MCP looks like. But it learns from what we inject into the system prompt, that the Sentry server is available to begin with. It's not completely guessing here.

      const orgs = await tools.mcp__sentry__find_organizations({});
      const { organizations } = orgs.structuredContent;
      const results = await Promise.allSettled(organizations.map(org =>
        tools.mcp__sentry__find_projects({
          organizationSlug: org.slug,
          regionUrl: org.regionUrl,
        })
      ));
      return organizations.map((org, i) => {
        const r = results[i];
        if (r.status !== "fulfilled") return { org: org.slug, error: String(r.reason) };
        if (r.value.isError) return { org: org.slug, error: r.value.content };
        return {
          org: org.slug,
          projects: r.value.structuredContent.projects.map(p => p.slug),
        };
      });
      

      Modern MCP Is A Fight

      I really don't want to talk too much about MCP here, but MCP is in fact a protocol that greatly benefits from Codemode. The problem in parts is that MCP in practice often targets harnesses that do not (yet?) use Codemode. But the tide is shifting. In the meantime, a temporary crutch has been to do what Cloudflare did, and do Codemode within the MCP server. But now we have Codemode in Codemode which is pretty bad. It means double JSON escaping, easy for smaller models to get confused by and the inner code cannot call the outer tools. So if you for instance use the Cloudflare MCP servers in Pi, the agent needs to write JavaScript and funnel it through more JavaScript. This is really not optimal, but it's also understandable that this is happening:

      const accRes = await tools.mcp__cloudflare__execute({
        code: `async () => {
          const r = await cloudflare.request({ method: "GET", path: "/accounts" });
          return r.result.map(a => ({ id: a.id, name: a.name }));
        }`,
      });
      const accounts = JSON.parse(accRes.content.map(c => c.text).join(""));
      
      const out = [];
      for (const account of accounts) {
        const r = await tools.mcp__cloudflare__execute({
          account_id: account.id,
          code: `async () => {
            const r = await cloudflare.request({
              method: "GET",
              path: \`/accounts/\${accountId}/workers/scripts\`,
            });
            return r.result.map(s => ({ id: s.id, modified: s.modified_on }));
          }`,
        });
        out.push({ account: account.name, workers: r.content.map(c => c.text).join("") });
      }
      return out;
      

      MCP Desires

      So to end things off: how well does Codemode work with MCP today? Well … not amazingly well. That's because MCP servers are not really targeting harnesses that use Codemode yet (though at this point I think most harnesses support it).

      For this to work well some recommendations:

      • Structured content: Codemode wants calls to return some nicely formatted JSON. So that needs to come back from the server, and many don't do that yet. The outputSchema system in MCP is great for that.
      • Consistent results: an interesting failure case is when an MCP server does not return consistent data. For instance because it tries to token optimize things depending on how many items are in the result set. This can cause an initial probe with 5 items to succeed, but then fail when the server returns the maximum batch size.
      • Large binary data: today MCP does not yet support large binary data so quite a few use cases that are really interesting do not work well at all yet. You end up with all kinds of weird workarounds such as pre-signed URLs to allow file uploads then to happen through non MCP channels.
      • Composable tool search: the MCP server might know better than the MCP client which tool is appropriate for a task. But there is no good mechanism today that allows a harness to fan out tool searches across multiple MCP servers. It's all emergent behavior and it does not scale well to multiple active servers.

      Future of Codemode

      So where does this leave us? Is this a reversal of what I wrote a year ago where I encouraged CLIs? I don't think so. In fact, the MCP ecosystem from my perspective picked up on exactly what we pointed out a year ago works: code. But Codemode goes beyond MCP in that it can act as a capable mechanism within the harness to express more freedom for the agent.

      There are however also some things that we still need to figure out. For one, durability with Codemode is trickier. We might have to adopt some ideas from durable workflow engines here to snapshot invocations. Or maybe, something like Starlark is a better composition language than JavaScript given its deterministic nature.

      Images, binary data and just the inability of this pattern to work with smaller models is also something that needs to be fleshed out. So it's for sure not a perfect solution yet, but it's quite a useful pattern that I expect us to leverage more.

    15. 🔗 Ampcode News Many, Many Pucks rss

      You can now have multiple separate conversations with Puck.

      We know you love Puck. We also know you've been asking Puck about your recent orbs, a bug fix, and dinner plans, all in the same conversation. Now each of those can be its own conversation, and each one gets its own Puck.

      The Puck window with the conversation menu open under the title. It lists five conversations, each with a differently colored Puck: Rebase vs merge, Watch the iOS release build, Draft Thursday's changelog, My review queue this week, and Ramen for six on Friday. Below them is an Inactive Last 72h group with two older conversations.

      Press + to start a new conversation. Click the conversation's name at the top to switch between them. Conversations you haven't touched in three days move into Inactive so the list stays short.

      Read more about Puck conversations in the docs.