- ↔
- →
- July 31, 2026
-
🔗 r/reverseengineering Akamai Bot Manager Sensor Data v3 — Fully Reversed & Deobfuscated 😳🤯 rss
submitted by /u/datadomee
[link] [comments] -
🔗 WerWolv/ImHex Nightly Builds release
-
- July 30, 2026
-
🔗 r/reverseengineering Patching my guitar amp's firmware rss
submitted by /u/igor_sk
[link] [comments] -
🔗 Hex-Rays Blog LLMs Have Reshaped How We Think About Decompilation and Collaboration rss
Shifting Times
A few weekends ago, I was playing DEF CON CTF Quals, the qualification event for the "Olympics of Hacking," with my team Shellphish. I say "playing" because I am, at this point, a washed-up hacker, but add me to the 40-something ranks! Nonetheless, I showed up, downloaded a binary, and went to open it in IDA Pro: a reflexive, built-in response to anything compiled. But as that ever-knowing face rendered on my screen, I paused for a moment to take in the surrounding feverish hacking. Something was strange: no one else had IDA Pro open. In fact, no one had any decompiler open!

-
🔗 hacker news ida pro references New comment by mahaloz in "Kuna: Decompiler Development in the Age of Coding Agents" rss
Interesting. A lot of data shows IDA Pro is significantly better than Ghidra: https://decbench.com/
Based on today's results, IDA Pro is ahead by 15 percentage points, which would mean, statistically, IDA Pro will recover perfect source code for 15% more functions than Ghidra on average.
-
🔗 hacker news idalib references New comment by billypilgrim in "Kuna: Decompiler Development in the Age of Coding Agents" rss
idalib with Claude Code already works really well. But honestly, despite what people have been saying, LLMs have been very good at decompiling for at least two years now, I have been using it for that purpose regularly. Even just copying disassembly from the current function and all nested called functions from IDA into ChatGPT is already unexpectedly good.
-
🔗 The Pragmatic Engineer The Pulse: Quitting Spotify Podcasts over reliability rss
Hi, this is Gergely with a bonus, free issue of the Pragmatic Engineer Newsletter. In every issue, I cover Big Tech and startups through the lens of senior engineers and engineering leaders. Today, we cover one out of four topics of last week 's The Pulse issue . Full subscribers received the article below seven days ago. If you 've been forwarded this email, you can subscribe here .
You can no longer watch The Pragmatic Engineer Podcast as video in the Spotify app (only as audio) because I have quit publishing video on that streaming platform. This comes after I decided that reliability takes a back seat within that team - and across much of Spotify. Unlike on other platforms such as YouTube, Apple Podcasts, and Substack, I've recently encountered a series of reliability issues around Spotify being unable to process video episodes. Even though I enjoyed a direct link with the Podcasts team there, things haven't improved.
So from now, I will no longer be publishing video episodes on Spotify. You can find videos of my in-depth chats with guests only on YouTube. Apologies for any inconvenience this change causes! Audio episodes of the podcast can still be found on Spotify via the RSS podcast feed hosted on Substack.
Honestly, the decision to quit the streaming giant wasn't hard, and I reckon there's a point here about the risk of deprioritizing reliable operations at major companies in order to push on things like AI adoption, as Spotify seems to be doing.
Some context: for the first two years of The Pragmatic Engineer Podcast, it was published on three podcast platforms:
- Substack 's podcast platform (audio): this is where the "master" RSS feed is served to the likes of Apple Podcasts, the web, Overcast, Pocket Casts, etc
- YouTube (video): video episodes uploaded individually
- Spotify (video + audio): every video episode was uploaded individually and then served as video or audio episodes from the platform.
As someone hosting a podcast, there are good reasons to bother doing three separate uploads:
- Most podcast platforms don 't support video. There will always be a need for a platform that serves the master RSS feed for audio versions while the video ones are elsewhere.
- YouTube doesn 't integrate with anything. YouTube is the leader in video podcast distribution, and uploading there directly makes sense.
- I had a direct line to the Spotify team, which was a big plus. Starting out the podcast, I had the unusual privilege of contact with the podcasts team, thanks to the newsletter gaining a decently-size audience. I was persuaded to take the plunge with them.
For eighteen months, nothing major went wrong. The admin portal for podcast publishers (called 'Spotify Creators') was pretty wonky; it gave intermittent errors, and was unable to remember me when I signed in, so, each Wednesday, I'd have to sign in with a code sent to my email to publish an episode.
But overall, things worked, until it all went suddenly downhill…
Unable to publish Spotify podcast episodes 3 out of 5 weeks
From late May, I did not include links to Spotify on new episode announcements because their podcasts product or platform seemingly had outages every time one published on Wednesdays at around 9am PST / 12pm EST / 6pm EU time.
Outage #1 (20 May): podcast publishing broke , my episode would not process on Spotify for 2+ hours. When uploading a video file to Spotify, there's a processing pipeline that runs to create chunks of the podcast in different video and audio formats. This pipeline appeared to stop running, meaning new episodes were not published.
It was not just the publishing that broke: the Creator portal looked absurd, with NaN% values everywhere, during the outage:
During
outage #1I emailed the Spotify team to alert them about the outage and also complained online. I got a response, confirming the outage and pledging to do better:
"The issue was in one of our podcast publishing metadata pipelines. A small subset of episodes completed normal media processing but then missed a downstream publish update because a newly introduced validation signal was not correctly wired into the logic that wakes up the publishing path. In simpler terms: the episode could become eligible to publish, but the final propagation step was not reliably triggered for that class of episodes.
We identified the root cause, deployed a fix, and reprocessed the affected episodes with all-clear called early this morning. We're also tightening the system so that fields used for publishing eligibility cannot be added without also triggering the relevant downstream updates.
Separately, we're reviewing how partial creator-impacting publishing delays are surfaced, because even when this is not a broad platform outage, it is still a bad experience for publishers like yourself.
Apologies again that you hit this. It was a real bug, not a wide outage, but it hit some of our most relevant creators."
Outage #2 (17 June): Spotify down. Four weeks later, when attempting to publish a video episode, all of Spotify went down for many users, including myself.
Spotify's
web player on 17 JuneSpotify does not maintain a status page, so it's impossible to tell how widespread the outage was. I didn't include a Spotify link in that week's announcement either.
Outage #3 (24 June): podcast publishing broke - again. Outage #3 in five weeks; deja vu. This time, it was episode publishing not working, yet again. After waiting two hours for the episode to publish on Spotify, I yet again sent out the announcement with no Spotify link.
I also emailed the Spotify Podcasts team, who confirmed the outage. I said I was considering stopping publishing video episodes, and to switch to audio- only publishing (which means pointing Spotify to my master RSS feed.) I said that an apology was appreciated but it wasn't enough to make it worth publishing video episodes there.
I also asked for the incident review because I had the feeling that reliability was not all that important on this podcast product. For the first outage I got a vague description of what happened, and promises of improvements that were never done - e.g. during this second outage, there was no improved communications to creators, which I was told would happen, after outage #1.
Internally, Spotify's team surely conducted an incident review as per usual, so I figured I'd hear back in about two weeks' time, and assumed a reply would be forthcoming because I'd made clear I was ready to leave Spotify Podcasts if reliability didn't improve.
No incident review three weeks later, so I quit Spotify
The incident review had never arrived as promised by three weeks later, even though there had been time for it to be completed. It was yet another sign of a platform that has become unreliable. Also, the creator portal occasionally threw up this error:
Spotify
's creator portal on 16 JulyI checked my Spotify stats: stream plays had been trending downwards unsurprisingly, given the ongoing outages, while the other podcast platforms didn't show the decline.**** It made me decide "enough is enough" and to move off Spotify.
Staying on their platform depended on seeing an incident review, but they didn't prioritize transparency, still had no status page, and nobody had built a feature for episode-processing status like YouTube has had for years. So, I pulled the plug and left:
Offboarding
from Spotify's (video) podcasts productAfter I made the switch away from Spotify, the platform's creators portal became buggier than ever, as in these examples:
My
Creators page after I changed the source of my podcasts to the master RSS feedComments disappeared:
My
show had no comments, suddenly… even though other parts of the UI showed dozens of comments:
Zero
comments, yet episodes with commentsEpisode links directed to 404 pages:
404
pages inside the Creator portal, when clicking linksA day or two later, these issues disappeared: I assume no one had tested the flow of moving away from Spotify Podcasts to an RSS feed, and it's why the experience was so poor.
Incident review finally published, but with a wrong timeline
A few days after offboarding from Spotify, their team published the incident report for outage #3. Reading through it, something did not add up in the timeline:
The
original timeline published for the 24 June incidentMy email account confirmed that I mailed the Spotify team at around 17:30 about the outage. So, after weeks of creating this report, why did the incident report downplay the fact that customers alerted the team before their own automated alerts fired?I complained to the Podcasts team, and to their credit, the incident report was updated:
The
updated incident timelineI didn't like how high-level the report is, and how vague the promised improvements were. Specifically, this one:
"During this incident, many creators learned something was wrong from their audiences before they heard anything from us. We are improving our processes and technical capabilities so creators get notified as soon as possible when things aren't working."
Overall, I don't regret the choice to leave, particularly when the focus of Spotify's leadership is on AI, not reliability.
Does Spotify have "AI psychosis?"
Previously, I used the term "AI psychosis" differently from the usual way of describing when someone starts believing everything an AI model tells them, however outlandish. I applied it to Meta's rush to develop its own AI model at the cost of the reliability of its profitable business activities. This was based on Instagram's most embarrassing-ever account takeover incident, which occurred when the team responsible for Instagram's Trust & Safety was slashed. Soon after, AI- generated, AI-reviewed code caused the hacking of a former US president's account.
At Spotify, it should have gone the other way. In March, I had the opportunity to meet its Head of Technology & Platforms, Tyson Singer, who said the company puts reliability far ahead of AI adoption, and doesn't adopt AI for its own sake. So, it was somewhat surprising to read the summary below of a podcast Spotify did with Anthropic:
"Spotify now ships 4,500 production deploys a day, and 73% of PRs are now AI-assisted.
Niklas Gustavsson (VP of Engineering at Spotify) keeps 5 to 10 Claude sessions running in tmux, one per git worktree, agents working in the background. All of it inside a 20M+ line monorepo. He expected agents to struggle at that size, but it's worked well.
Spotify's migration codemods grew into thousands of lines of edge cases. Code has too much API surface for static rewrites. Early LLMs barely did better. Adding a judge took PR success from ~25% to 80%.
All of this leans on verification, the single most important thing when agents are used and the place most companies underinvest
Spotify rebuilt their test automation around it so engineers can confidently guide and supervise agents, rather than manually execute repetitive tasks."
It seems to me that all the talk is about usage of AI, and none about reliability , all while Spotify's platform becomes less reliable than ever, at the same time as the streamer is going all-in on AI; with AI judges and devs running 5-10 parallel Claude sessions.
All things considered, it's worth asking if Spotify has the corporate variant of "AI psychosis", whereby the reliability of a successful operation gets torched in the chase for the next big thing by executives. I don't even think Spotify is all that different from Meta and other companies in this!
Things look bad, based on the quality and reliability degradation of products. Annoyingly, in many cases, customers don't really have the choice of going elsewhere. My podcast is an exception, as video podcasts on Spotify never truly took off, so quitting the platform wasn't a big deal. Even so, I'm particularly disappointed that Spotify has prioritized AI usage over reliability. I know some executives there pushed against this, but I feel safe in assuming that they lost that battle.
Value of staying reliable & "sucking less"
Max Kanat-Alexander, distinguished engineer at Capital One, has written about how a software project can become wildly successful just by "sucking less" in his reflections upon the success of the Bugzilla project, (2004-2009):
"All you have to do to succeed in software is to consistently suck less with every release.
Nobody would say that Bugzilla 2.18 was awesome, but everybody would say that it sucked less than Bugzilla 2.16 did. Bugzilla 2.20 wasn't perfect, but without a doubt, it sucked less than Bugzilla 2.18. And then Bugzilla 3.0 fixed a whole lot of sucking in Bugzilla, and it got a whole lot more downloads.
Why is it that this worked?
As long as you consistently suck less with every release, you will retain most of your users. You're fixing the things that bother them, so there's no reason for them to switch away. Even if you didn't fix everything in this release, if you sucked less, your users will have faith that eventually, the things that bother them will be fixed. New users will find your software, and they'll stick with it too. And in this way, your user count will increase steadily over time.
But what happens if you release frequently, but instead of fixing the things in your software that suck, you just add new features that don 't fix the sucking? Well, eventually the patience of the individual user is going to run out. They're not going to wait forever for your software to stop sucking."
Personally, I got tired of Spotify's Podcasts product continually going in the wrong direction on Max's scale: the poor reliability, frequent errors on the Creators site, and the sense that they don't really care about improving existing things.
Read the full issue of last week's The Pulse. The full The Pulse additionally covers:
- Will Kimi K3 trigger US push for closed-source AI models? Moonshot AI's latest open model, Kimi K3, is on par with Anthropic's Fable 5. Could it lead to the US government regulating or banning Chinese open models to protect US labs?
- AWS laughs off "heart attack" billing error. AWS customers were billed trillions more than they should have been, due to what was likely a conversion error. But instead of sharing an incident report, AWS saw the funny side.
- Industry pulse. OpenAI's unreleased model tried to hack HuggingFace to improve its test scores, X took more than a year to develop its new Android app, Google's new AI model flops, and more.
-
🔗 r/reverseengineering Reversing of Eufy Security Video Doorbell sync protocol and wifi creds decryption from flash memory rss
submitted by /u/gid0rah
[link] [comments] -
🔗 hacker news ida pro references New comment by comandillos in "Kuna: Decompiler Development in the Age of Coding Agents" rss
Haven't used IDA much lately, but after looking at the screenshot with that IDA PRO decompiled code in their website I feel like Ghidra is already ahead of them in this area :D
-
🔗 r/reverseengineering I reverse-engineered Intel's HECI protocol and built a Python tool that talks to the ME directly rss
submitted by /u/Frequent-Ad-9633
[link] [comments] -
🔗 Kagi release notes July 30th, 2026 - Kagi Assistant on the go and design refinements for Search rss
Announcing the official Kagi Assistant apps
Kagi Assistant is now available as a native app for iOS and Android!
Ask a question, explore the web, work with files, conduct in-depth research, or choose from leading AI models, all from your phone. Your threads and Custom Assistants stay with you, so you can pick up wherever you left off.
These are the first steps towards delivering a fantastic Kagi Assistant experience on mobile, with much more to come.
Download it now:
- App store: https://apps.apple.com/app/6755965340
- Play store: https://play.google.com/store/apps/details?id=com.kagi.assistant
Give it a spin and let us know what you think!
Report responses directly from Kagi Assistant
You can now report an assistant response without leaving the conversation. Hover over any assistant message and select the thumbs-down button to open the feedback form, where you can report issues for reasons ranging from UI bugs to harmful content.
Note that when you submit a report, the full thread is shared with Kagi for review. The report and its associated copy of the thread are automatically deleted from Kagi’s review records after 30 days.
Export or delete all your threads
We've also added important controls, so you can now export all your threads or permanently delete them at once from
Settings > General.
Kagi Search
A sharper search experience
We’ve polished the search results page to make its controls easier to find and understand. From the filter bar to domain-related options and menus, these updates bring greater clarity and ease of use to the features you rely on most.

Exchange rates, right in your search results
Next up in our broader effort to improve search widgets: currency conversion. Comes handy when you’re planning a trip, shopping abroad, or simply want to keep tabs on exchange rates.

Other improvements and bug fixes
Kagi Search
- Fixed several animations that didn't respect the system's
prefers-reduced-motionsetting - 'Open first result' shortcut suggestion tries to escape double-quotes #8752 @craftypersimmon
- Fake 1337x domain #9279 @fxgn
- Dice Number getting cut off in the thousands #10964 @Flossiii
- Some Kagi lenses not working for me in Kagi search #10970 @Fearce
- NSFW results when searching for "xteink black vs white" while safe search is turned on #10905 @ciccero040
- Incorrect definition of "socialism" #10988 @thoroughly
- Nothing triggers the weather widget when the interface language is set to German #4612 @laiz
- Cannot Manually Select Location in Privacy Settings #11005 @iamjameswalters
- More and share buttons disappeared - Mobile DOM #10957 @NyraSyn
- Slopstop blocks whole domains #11039 @kslays
- Unable to report AI image slop on mobile due to popup closure #11066 @Hanbyeol
- Kagi adding extra {{{s}}} in bang redirect when no query #9885 @jadams9
- "Profile not found for ki_research" error when using
??shorthand #11020 @paying_customer - Kagi Knowledge answer for "Labour Day 2025" gives wrong date #8713 @wanion
- Delete recent language option #8549 @ten
- Stopwatch should not start from searches like
0424:2422#6061 @xfhrnozxqnrnqrsvntp - [Android] Quick Switch Doesn't work #7695 @cr0ntab
- Save a round trip: Advertise HTTP/3 support in an HTTPS DNS record #10829 @drrlvn
- Searching for
<script>returns no results #11117 @Bonarc
Kagi Assistant
- Camera button in assistant #5261 @Arnaud
- Assistant error "something went wrong..." when "web access" is selected #6687 @Nyaa
- Gemma 4 31B failing to read images #10947 @Dustin
- Assistant: Remove whole history #6971 @Wanja
- Didn't like a Quick Answer response? post it here #9082 @Thibaultmol
- Hourglass design is bad #11034 @shurik
- FIXED - !ai bang - Query not working #11077 @fanged_bagful
- Temporary threads don't actually get removed #10780 @foxberg34
Kagi Translate
- Kagi Translate reloads the page when using website translate #10852 @tijol
- Dictionary now shows language-specific grammar details, starting with Czech animate/inanimate nouns
- Proofread no longer suggests changes to text that was already correct, such as de-capitalizing German nouns
- Translations no longer occasionally come back untranslated
- Translations keep proper typographic punctuation instead of straightened quotes
- Alternative translations now work when selecting part of a longer text
- Double and triple-click selection in the translated text works as expected, and the alternatives panel no longer flickers while loading
- "New version available" banner appears less often and supports dark mode
- Reset-All Button for Translate #10999 @erakagi
- Document Translate for Typst #11098 @weriomat
- Palestinian Arabic in Translate #11006 @zsoltsb
- Phonetic Translation Placement #10699 @dwahdany
- Myanmar alias for Burmese #10495 @mb
- Translation History panel cannot be closed in Brave (Windows 11) #10997 @vshlapakov
- Prompt being read prior to translated word #10974 @kagifeedback-1xxkg
- Kagi Translate Audio Broken #10923 @levers
Kagi News
- Kagi News no longer showing births and deaths Today In History #11026 @eggnog_90241
- Content filter should backfill removed articles with the next available ones #10828 @roy_carver023
- Clearer feedback door #10919 @Coldcartcold
- Fast font option #10001 @broken665
- RTL content is displayed in the wrong direction when the interface language is LTR #10298 @maxmellen
-
🔗 Cryptography & Security Newsletter The State of Post-Quantum Cryptography rss
For people who deal with network security, the last couple of years have been busier than usual, largely because of the impending threat of quantum computers that are going to annihilate all cryptography we use and love today. If you’re feeling fatigue from the firehose of events, you’re not alone. For me, personally, there is a sense that we’re seeing fewer and fewer technical announcements, which is perhaps pointing to the fact that most of the technical decisions have been made. At the same time, the deadlines have been shortened on account of the fears that cryptographically-relevant quantum computers (CRQC) are arriving sooner than anticipated. What lies ahead of us is definitely going to be hard, but we’ll need to only execute on the plans?
-
🔗 r/reverseengineering Kuna: Decompiler Development in the Age of Coding Agents rss
submitted by /u/mttd
[link] [comments] -
🔗 r/reverseengineering Zion Basque releases Kuna, a decompiler built almost entirely by an LLM rss
submitted by /u/ryanmerket
[link] [comments] -
🔗 pydantic/pydantic-ai-harness v0.14.0 (2026-07-29) release
What's Changed
- fix(ci): let dependency approval revoke a label and still report by @dsfaccini in #495
- docs(agent_docs): record where operational policy lives by @dsfaccini in #497
- Clarify runtime capability creation documentation by @dsfaccini in #499
- Add first-party MongoDB backend + externalize large text parts by @dsfaccini in #446
Full Changelog :
v0.13.0...v0.14.0 -
🔗 openonion/connectonion Release v1.5.2 release
Two things that were silently broken.
🐛 Fixes
Claude calls were dropping the system prompt entirely
(#289)
Anthropic takes the system prompt as a top-level
systemargument rather than a message, so_convert_messagescorrectly removed system messages from the list — with a comment saying they would be "handled separately". That separate handling was never written. Neithercomplete()norstructured_complete()ever set it.The result: every agent on a
claude-*model ran with no system prompt at all — no persona, no instructions, no constraints. No error, no warning. If you use Claude models, take this release.Fixed in #293 by @sanmaxdev.
.co/docs/was empty on every PyPI install(#267)
co inittells you ".co/docs/for full documentation". The folder was empty.docs/lives at the repo root, and the wheel is built frompackages = ["connectonion"], so the documentation was never inside the package — the sdist had it, the wheel did not, andpip installuses the wheel.copy_docs()printed a yellow "Documentation not found" warning and left you an empty directory.The files now ship inside the wheel. After
co inityou get 168 markdown files in.co/docs/— verified by installing from PyPI into a clean venv and runningco init.Installation
pip install --upgrade connectonionBreaking Changes
None.
Full Changelog :
v1.5.1...v1.5.2 -
🔗 Console.dev newsletter superfile rss
Description: Fancy modern terminal file manager.
What we like: Built with Go. Supports all the file operations you’d expect. Build-in search. Split into panels and copy/paste between them with shortcuts. Configurable and themeable.
What we dislike: Partial Windows support.
-
🔗 Console.dev newsletter LetsSeal rss
Description: Prove files.
What we like: Built on an open standard to prove a file exists and is unaltered. Works via the web and through a CLI. Has a GitHub action. Can be self-hosted. Verification is through re-hashing and comparison, with a public transparency log.
What we dislike: Although you can verify offline, it still requires writing into a public transparency log (blockchain).
-
- July 29, 2026
-
🔗 IDA Plugin Updates IDA Plugin Updates on 2026-07-29 rss
IDA Plugin Updates on 2026-07-29
Activity:
- diaphora
- 842d83ba: Merge pull request #363 from r0ny123/fix-plugin-logger-isolation
- haruspex
- hrtng
- 6bfe8fb6: Enums: fix issues raised in #60
- ida-pro-mcp
- IDAPluginList
- 2411ad74: chore: Auto update IDA plugins (Updated: 19, Cloned: 0, Failed: 0)
- Luc-Nhan
- mcrit-plugin
- aafb98c2: Fix generator len() error in GUI smoke block widget exercise
- 1e41c767: Suppress IDAPython PyQt5 shim dialog in GUI smoke
- feab22fb: Harden IDA GUI smoke startup
- 9b4e85d4: Align IDA GUI smoke Python runtime
- 21c4210f: Fix GUI IDA startup configuration in CI
- 638837cf: incremental fixes part 23
- 499afcbf: incremental fixes part 22
- c852e9f1: incremental fixes part 21
- b0e8af12: incremental fixes part 20
- 2b2094d1: incremental fixes part 19
- 9abd668b: incremental fixes part 18
- edbd3845: incremental fixes part 17
- c3eb04cd: incremental fixes part 16
- cde467c8: incremental fixes part 15
- 8c1a167e: incremental fixes part 14
- c2894b05: incremental fixes part 13
- ffa3b884: incremental fixes part 12
- bf25ee68: incremental fixes part 11
- b6ac0c88: incremental fixes part 10
- 66f7db90: incremental fixes part 9
- pharos
- rhabdomancer
- 54fcba5e: refactor: optimize performance in the implementation of
Priority
- 54fcba5e: refactor: optimize performance in the implementation of
- diaphora
-
🔗 r/reverseengineering GitHub - SarangRao20/battery-charge-limiter: Reverse-engineered the Embedded Controller register for battery charge control. Built a cross-platform daemon (Windows + Arch) that enforces 80% hardware cap on laptops where BIOS hides this feature. Includes EC discovery tool for other models. rss
submitted by /u/NobitaNobi12345
[link] [comments] -
🔗 earendil-works/pi v0.83.0 release
New Features
- Credential export for external clients —
pi auth print-api-keyandpi auth print-bearer-tokenexport configured credentials with automatic OAuth refresh and minimum-validity enforcement. - Headless OpenRouter sign-in — Complete
/loginover SSH by pasting the redirect URL or authorization code when the loopback callback is unavailable. See OpenRouter. - Claude Opus 5 on GitHub Copilot — Use Claude Opus 5 through GitHub Copilot with adaptive thinking and a 1M context window. See GitHub Copilot.
Breaking Changes
- Upgraded bundled TypeBox aliases to 1.3.7, removing deprecated APIs including
Type.Base,Type.Awaited,Type.Promise,Type.AsyncIterator,Type.Iterator,Type.Options, andValue.Mutate, while fixing compiled validation of nullable array tool arguments. Extensions using removed APIs must migrate to supported TypeBox APIs. See Package Dependencies (#7243 by @petrroll).
Added
- Added
pi auth print-api-keyandpi auth print-bearer-tokencommands for exporting configured credentials to external clients, including automatic OAuth refresh and configurable minimum token validity (#7168). - Exposed the session's resolved model scope as
ctx.scopedModelsto extensions. See Extension Context (#7191 by @pungggi, #7215). - Added inherited per-request
fetchinjection for supported text and image provider transports. - Added the inherited
"pending"stop reason for partial streaming messages. See Custom Provider Stream Pattern (#7151 by @lucasmeijer). - Added inherited raw provider stop reasons across Google, Anthropic, Amazon Bedrock, Mistral, and OpenAI streams; unmapped terminal reasons now surface as provider errors instead of successful stops (#7272).
- Added manual redirect URL and authorization-code entry to OpenRouter login for remote and headless environments. See OpenRouter (#7114 by @rgarcia).
- Added inherited Claude Opus 5 support for GitHub Copilot with adaptive thinking and a 1M context window. See GitHub Copilot (#7158 by @jay-aye-see-kay).
Changed
- Changed inherited OAuth credential resolution to refresh tokens with less than five minutes of validity remaining instead of waiting until expiration (#7168).
Fixed
- Added a status line when the tool output expansion is toggled (#7180).
- Fixed file-backed
SYSTEM.mdandAPPEND_SYSTEM.mdprompts being omitted from the interactive startup context listing. See System Prompt Files (#7096). - Fixed context files loading twice when a linked Git worktree is nested under its main repository. See Context Files (#7221 by @arajkumar).
- Fixed llama.cpp streamed responses reporting zero token usage and leaving session context accounting empty. See llama.cpp (#7258 by @SteveImmanuel).
- Fixed session replacement and committed tree navigation during an active response to abort and persist the outgoing turn instead of leaving dangling tool calls. See Sessions (#7022 by @tmustier).
- Fixed failed Git package installs leaving partial directories that blocked clean retries. See Install and Manage (#7210 by @haoqixu).
- Fixed the
/modelselector retaining a stale selection while filtering instead of highlighting the top match (#7211 by @christianbasch). - Fixed direct RPC bash commands bypassing extension
user_bashhandlers. See User Bash Events (#7214). - Fixed skills, prompts, and themes losing package source metadata after extensions reload resources. See Resource Events (#6968).
- Fixed cancellation of concurrently running user bash commands so every active command is aborted (#7103 by @yzhg1983).
- Fixed duplicate messages appearing when extensions switch sessions during interactive startup (#7110 by @yzhg1983).
- Fixed inherited Qwen Token Plan reasoning models to send their service-specific thinking controls and supported reasoning-effort levels (#6951, #6998).
- Fixed inherited Z.AI output limits being sent through an unsupported parameter. See Providers (#7174 by @HyeokjaeLee).
- Fixed explicitly configured Amazon Bedrock profiles being overridden by ambient AWS access keys. See Amazon Bedrock (#7176 by @christianbasch).
- Fixed inherited image fallback paths overflowing narrow terminals, shortened home-directory paths, and made absolute paths clickable when terminal hyperlinks are available (#7262).
- Fixed inherited OpenAI-compatible tool calls losing their function arguments when malformed deltas also contain an empty
customobject (#7288 by @sunnyyoung).
- Credential export for external clients —
-
🔗 tomasz-tomczyk/crit v0.18.2 release
What's Changed
Security
- fix: require explicit ack for unauthenticated network exposure (#776) by @tomasz-tomczyk in #776
- fix: reject cross-site browser POSTs via Sec-Fetch-Site (#775) by @tomasz-tomczyk in #775
Agent integrations
- feat: add ampcode integration for crit install (#761) by @mochadwi in #761 - Thank you!
- feat: configure Claude plan approval mode (#759) by @tomasz-tomczyk in #759 - Thank you @keoz-higidi for suggesting!
- feat: opt-in round-ready notify and OpenCode wait toast (#755) by @tomasz-tomczyk in #755 - Thank you @FlorentDhamma for suggesting!
- fix: correct opencode sub-command from 'ask' to 'run' (#754) by @blyoa in #754 - Thank you!
- fix: require explicit Crit invocation across agent integrations (#756) by @tomasz-tomczyk in #756 - Thank you @rsanheim for raising!
Notifications & review UX
- feat: auto-close review tab after approval (#753) by @tomasz-tomczyk in #753 - Thank you @jvaldiviezo9 for suggesting!
- feat: treat --output as crit data root for keyed reviews (#771) by @tomasz-tomczyk in #771
- fix(cli): add command-scoped help (#770) by @tomasz-tomczyk in #770 - Thank you @pstibrany for suggesting!
- fix: clear auto-close on dismiss and drop inert browser notifications (#773) by @tomasz-tomczyk in #773
- fix: avoid duplicate browser tab on live cold start (#768) by @tomasz-tomczyk in #768 - Thank you @CarlosZ for reporting!
- fix: honor configured output across review commands (#769) by @tomasz-tomczyk in #769 - Thank you @CarlosZ for reporting!
- fix: honor legacy --output reviews and close browser-open gaps (#772) by @tomasz-tomczyk in #772
Dependencies
- chore(deps-dev): bump eslint from 10.6.0 to 10.7.0 (#751) by @app/dependabot in #751
- chore(deps): bump actions/setup-go from 6 to 7 (#749) by @app/dependabot in #749
- chore(deps): bump actions/setup-node from 6 to 7 (#750) by @app/dependabot in #750
New Contributors
Full Changelog :
v0.18.1...v0.18.2 -
🔗 @HexRaysSA@infosec.exchange Vegas in August = Hacker Summer Camp. mastodon
Vegas in August = Hacker Summer Camp.
We'll be at Black Hat, B-Sides LV, and DEF CON 34.Highlights: a demo of our upcoming Malware Analysis Add-On, a hands-on DLL sideloading workshop (40 seats, register now), a live look at Teams' new Git- native workflow, and recruiting. (Yes, we're hiring!)
👉 Find the full rundown and where to catch us: https://hex-rays.com/blog/hex- rays-hacker-summer-camp-2026
-
🔗 r/reverseengineering pwnable.kr - mistake rss
submitted by /u/AdvisorPowerful9769
[link] [comments] -
🔗 openonion/connectonion Release v1.5.1 release
What's Changed
✨ Improvements
co statusnow says where every API key comes from — process environment,<project>/.env, or~/.co/keys.env— and flags keys defined in more than one place as a conflict, so a stale duplicate stops being invisible. Add--revealto print the values. (#248)
📚 Housekeeping
- The project states Apache-2.0 everywhere, matching the LICENSE file. README and
pyproject.tomlpreviously said MIT. (#291)
Installation
pip install connectonion==1.5.1Breaking Changes
None.
Full Changelog :
v1.5.0...v1.5.1 -
🔗 @binaryninja@infosec.exchange Current Binary Ninja newsletter subscribers are automatically entered. New mastodon
Current Binary Ninja newsletter subscribers are automatically entered. New subscribers who sign up during the giveaway will also be entered for remaining drawings. Sign up here: https://v35.us/dn6rcg5
-
🔗 @binaryninja@infosec.exchange Day 7 of our 10-year anniversary celebration comes with another big prize! mastodon
Day 7 of our 10-year anniversary celebration comes with another big prize! Today we’re giving away a Commercial license! Already have a license? The prize can be used as a license extension. https://binary.ninja/10years
-
🔗 r/reverseengineering Hi everyone, I've released APKX-Hunter v2.5.0, an open-source Android Static Analysis Framework written in C. rss
submitted by /u/SyscallX-18113
[link] [comments] -
🔗 r/reverseengineering Apple APTicket / LocalPolicy Forensic Kit rss
submitted by /u/penwellr
[link] [comments] -
🔗 pydantic/pydantic-ai-harness v0.13.0 (2026-07-28) release
What's Changed
- subagents: per-delegation model selection via an opt-in model menu by @dsfaccini in #451
- Register skills page in docs nav.json by @dsfaccini in #484
- Rename capabilities to follow the documented naming convention by @DouweM in #480
- Skip LocalStack integration on fork pull requests by @dsfaccini in #487
- docs(agents): push-and-watch rule for agent contributors by @dsfaccini in #452
- Keep pre-rename PyaiDocs agent specs loading by @DouweM in #488
step_persistence: opt-inmax_snapshots_per_runto bound snapshot growth by @dsfaccini in #442- test(docs): enforce
docs/nav.jsonparity so pages cannot ship orphaned from the site nav by @dsfaccini in #491 - Gate dependency-file changes on maintainer approval by @dsfaccini in #492
- feat(
conversation_search): BM25 search over the historyStepPersistencestores by @dsfaccini in #413
Full Changelog :
v0.12.0...v0.13.0 -
🔗 seanmonstar Micro: I want your own words rss
If you choose to communicate with me, all I ask is that you use your own words. Bug reports and issues. Pull request descriptions and especially review comments. This isn’t new, but I wanted a link of my own.
LLMs write way way too much. I don’t know if you understood it enough for me to ask questions back. They don’t use reasoning, so is it real? If you didn’t care enough to write it, do I care to read it?
I’m torn when I receive LLMed reports. Initially, I really don’t want to read all that. But at the same time, my mind nags me; something’s broken and I should fix it for everyone. If it makes it easier for someone to report a bug, that’s a pro, I guess.
-
🔗 Mitchell Hashimoto Superlogical rss
(empty) -
🔗 Ampcode News Banking on the Frontier rss
Amp Labs has partnered with Westpac, one of Australia's oldest and most respected companies, a fixture of Australian business for more than two centuries, to transform how the bank builds and delivers technology.
Amp Labs operates on the thesis that the biggest impact comes from working with carefully selected customers with deep mutual trust and a shared ambition to explore the frontier together.
The frontier isn't just the models and the agent, it's putting them to work on real enterprise problems, like migrating and modernizing data systems that millions of people rely on every day. Being on the front lines and doing the work inside these companies is how we keep Amp on the frontier for everyone.
We are building a team to work side-by-side with Westpac engineers, on site in Sydney.
The Founding Team
Gareth Townsend Previously Block
Matty Evans Previously Ethereum Foundation
Andrew Gerrand Previously Google
Chris Nicol Previously Canva
Ryan Christensen Previously CanvaJoin Us
We're hiring. If you are Sydney-based and want to work on the frontier inside Australia's first company: join@amplabs.com.
-
🔗 Ampcode News Who Cares About the Model? rss
Two weeks ago we shipped the Dial and quietly did something that's supposed to be traumatic: we changed the default model.
Before the Dial, Amp's default mode was
smart, running Claude Opus 4.8 and carrying more than half of all new threads. The Dial mademediumthe default, andmediumruns GPT-5.6 Sol.Most Amp users switched from Anthropic to OpenAI overnight.
We braced for the outcry. Every model swap in every coding tool comes with one. We prepared migration docs, packaged the old modes as installable plugins, and waited.
Nothing happened. Not a single complaint.
Here's what the switch looked like in production:
The day before the Dial shipped,
smartcarried 55% of new threads. A week later: zero. Last week, the four Dial modes carried 93% of all new threads, andmediumalone carried two-thirds. Of the users on the Dial, 69% never set it to anything butmedium.And the Amp plugins we shipped to bring the old modes back — exact prompts, exact models, one command to get
smartagain? Almost nobody installed them.What This Tells Us
The differences between frontier models are now small. Small enough that for one engineer on one task, switching models won't visibly change the result.
But they still matter to us, here at Amp: across every thread on every tier, small differences compound, so we keep benchmarking and swapping models. That's the trade. The default is good because someone is paid to care about it, and it doesn't have to be you.
What does change your output are three things, none of which is a model: how hard the task is, what context you put in, and how closely you review what comes out. All three have more impact on the outcome than whether you use this or that latest frontier model.
The Dial covers the first one: how hard the task is. You tell it how hard the task is and the corresponding setting on the Dial uses whichever model wins it right now, re-tested constantly. You know your tasks, we know the models.
And for the other two — the context and the review — the rest of Amp lets you do the best job with that.
-
- July 28, 2026
-
🔗 IDA Plugin Updates IDA Plugin Updates on 2026-07-28 rss
IDA Plugin Updates on 2026-07-28
Activity:
- augur
- 06b349c5: doc: update CLAUDE.md
- capa
- CTFStuff
- haruspex
- hexmux
- ida-domain
- 883ae251: Include examples in package
- leaknet
- 9adecb6f: -many fixes:
- mcrit-plugin
- quokka
- 7ecc172e: Merge pull request #135 from quarkslab/dependabot/github_actions/acti…
- twdll
- zydisinfo
- 02a7f1b5: Update CMake version to 4.2.x and modify generator arguments for Wind…
- 5dcbe53a: Add support for Visual Studio 2022 generator in CMake build step on W…
- 619d8cdd: Update MSBuild setup to use version 2 and remove Visual Studio 2022 s…
- 8127da2d: Update event handling to use ui_screen_ea_changed for address updates
- augur
-
🔗 Hex-Rays Blog Hex-Rays is Heading to Hacker Summer Camp 2026 rss
-
🔗 r/reverseengineering Reproducing a D-Link firmware CVE on an emulated FirmAE twin, packaged as a signed receipt you can re-verify in 5 min (with a PR:N→PR:L CVSS correction) rss
submitted by /u/TheLuisBolivar
[link] [comments] -
🔗 r/reverseengineering The Elevator Glitch: How One Function Destroyed Public Lobbies in CoD4 rss
submitted by /u/Rex109
[link] [comments] -
🔗 r/reverseengineering Phantom Stealer – jsc.exe Injection, Credit Card Theft & Email Exfiltration rss
submitted by /u/StructBreaker
[link] [comments] -
🔗 @binaryninja@infosec.exchange Current Binary Ninja newsletter subscribers are automatically entered. New mastodon
Current Binary Ninja newsletter subscribers are automatically entered. New subscribers who sign up during the giveaway will also be entered for remaining drawings. Sign up here: https://v35.us/dn6rcg5
-
🔗 @binaryninja@infosec.exchange Today’s giveaway is huge! It includes one year of Binary Ninja Non-Commercial mastodon
Today’s giveaway is huge! It includes one year of Binary Ninja Non-Commercial PLUS one year of Sidekick Non-Commercial! Already have a license? The prize can be used as a license extension. https://binary.ninja/10years
-
🔗 r/reverseengineering I reverse engineered an ASUS embedded controller's fan protocol. rss
submitted by /u/Keyitdev
[link] [comments] -
🔗 r/reverseengineering Reverse-engineered the BLE protocol of a discontinued Fisher-Price toy Lumalou after its app was discontinued rss
submitted by /u/EmanueleStrazzullo
[link] [comments] -
🔗 HexRaysSA/plugin-repository commits sync repo: +1 release rss
sync repo: +1 release ## New releases - [rhabdomancer](https://github.com/0xdea/rhabdomancer): 0.10.0 -
🔗 pydantic/pydantic-ai-harness v0.12.0 (2026-07-27) release
What's Changed
- Add Agent Skills as deferred capabilities by @adtyavrdhn in #396
- Warn on overlong Agent Skill descriptions by @adtyavrdhn in #466
- Bump Pydantic AI to 2.18.0 by @adtyavrdhn in #472
Full Changelog :
v0.11.0...v0.12.0 -
🔗 r/reverseengineering GitHub - memues/mstar-monitor-firmware-dumper: Dump SPI flash firmware from MStar scaler-based monitors over I2C/DDC rss
submitted by /u/AtheistMonkeys
[link] [comments] -
🔗 openonion/connectonion Release v1.5.0 release
[1.5.0] - 2026-07-28
✨ Features
- Agent Home pages. An agent can keep a
dashboard.htmlin its project root; the host reads it and pushes it over the existing agent WebSocket, so a chat client can render it beside the conversation. Sent on connect and after any run that changed the file, with a per-connection(mtime, size)check so an untouched Home costs nothing per turn.host()writes a polished starter on day zero and never clobbers yours. Buttons carryingdata-ochat-skillrun a skill as a visible chat turn. See docs/network/dashboard.md. co syno— drive your Synology NAS from the terminal.co aiYOLO mode — skip the approval prompts when you want it to just go.- Agents can declare how long they need a browser tab , so a long job isn't cut off by tab contention.
♻️ Changed
co aiand the project templates drive the browser throughco browserrather than 40 in-process tools — one browser story instead of two.- Deploy polls the full build window and validates the project name locally before uploading, so a typo fails fast instead of halfway through a build.
🐛 Bug Fixes
- The
[env]diagnostic only prints when stderr is a terminal, so it no longer pollutes piped output. - A live daemon's socket is no longer unlinked on a non-refusal
OSError. - Unit tests are hermetic, and
send_emailworks without a.envfile.
[1.2.1] - 2026-07-17
✨ Features
co browserruns natively on Windows — the daemon speaks named pipes (stdlib, HMAC-authenticated) on Windows and keeps its Unix socket byte-identical on macOS/Linux. No WSL.- Zero-setup browser : the first page-driving
co browsercommand auto-installs chromium (per-user, no admin rights) when no browser exists; desktop Chrome is auto-detected and preferred.
🐛 Bug Fixes
- Windows: emoji/Unicode CLI output no longer crashes on legacy codepages (cp1252) — including when
cois driven through a pipe by Claude Code/codex. co browser go_tono longer rewritesfile:///about:/data:URLs into brokenhttps://forms.- Offline first run:
co initnow degrades gracefully (friendly one-liner, scaffold + keypair still created) instead of aborting with a traceback;authenticate()carries a 15s network timeout. - Docs: removed nonexistent
co init --no-ai/--update-docsflags; fixed the stale auth URL reference.
✅ Testing
- New
windows-e2eCI job simulates a real Windows user end to end: wheel install, PowerShell 7 / 5.1 / cmd.exe, offline first run, CJK + space home directory, fresh-laptop browser auto-install (real download), three parallel agents, Task-Manager-kill recovery, authkey self-healing. - Windows unit/daemon matrix on Python 3.10/3.12/3.13; concurrency tests (8 simultaneous clients, cold-start singleton race).
[1.0.5] - 2026-07-01
🐛 Bug Fixes
co email:get_emailsnow reads the API'stext/htmlbody fields (falling back to legacytext_body/html_body), soco email inboxandco email readshow the message body instead of an empty string.co browser:take_screenshotprints the saved file path instead of dumping the base64 image to the terminal.
📚 Documentation
- Agent Home pages. An agent can keep a
-
🔗 Probably Dance If AI Writes All the Code, What Do the Programmers Do? rss
Eight months ago I was producing roughly 90% human written code and 10% AI written code. This has switched surprisingly rapidly and now my code is probably 90% AI. So what do I do all day?
I'll go through a change that I made to a matrix-multiply kernel, for which I'll sadly have to be a bit vague. Since GPUs are now giant matrix-multiply chips, where north of 96% of the flops are in the tensor cores (latest Nvidia GPUs have 2250 tflops in bfloat16 matmuls, compared to 75 tflops for everything that's not a matmul), you'd think that they'd make it easy to use all those flops. But no, matrix multiply kernels are giant crazy beasts that are incredibly tricky to get right. Some quick googling finds this explanation on Nvidia hardware and this one on AMD hardware. Just open those two and scroll down both to get a feeling for how much work is involved. (and yes, these implement the simple O(n^3) matmul loop where you iterate along repeatedly and multiply and add all the numbers)
The particular kernel I was optimizing was using Cluster Launch Control (CLC), a complexity that's covered in neither of the blog posts linked above, and it had some bad interactions with some particular inputs. The resulting change was 95% written by AI. So I'll just go through what I did. AI asks are bolded:
0. Find and Understand the Problem
Before I started I had to find and understand the problem. Some coworkers had talked about the matmuls taking too long, and nsys showed that the tensor- cores were oddly idle while that matmul kernel was running, and then it took some more targeted benchmarking to confirm that slightly different inputs give big speedups, which convinced me that something silly must be going on. My coworkers actually had a really good theory, which brings me to the actual start of the work on this task:
1. Find the Code
The code was in a library that I was vaguely familiar with, but since matmul kernels are giant scary beasts, it was a bit hard to navigate. So before I even tried, I asked an AI to find the relevant code for me. I explain what the problem is and I explain our theory. Then I also start looking myself but the AI finds the relevant code first and even points me at exactly the right lines to look at. It also confirms that our theory for the problem sounds plausible.
2. Understand the Code
Next I decide to understand the code myself, so I intentionally don't ask the AI anything more. As I look around I begin to think that our theory actually isn't right. I mean it was partially right, but the code very much wasn't doing the slow thing that we thought it was doing. It was doing something slower: returning out of CLC mode and giving control back to the hardware scheduler.
This is a little silly, so after I am convinced of it, I ask the AI to confirm , just to get a second opinion. It agrees with me (but I take that with a grain of salt, because it also agreed with the initial theory).
3. Monkey-patch the Code
Since neural network training happens in Python, you fix things by monkey- patching. At least initially this is the fastest way to iterate on this code. But monkey-patching has a bit of boiler-plate that's easy to get wrong, so I ask the AI to set it up for me.
Then I try to fix the bad behavior, but my fix results in a deadlock. I look at the surrounding code again, but realize I don't understand CLC enough (this is my first interaction with it, and I jumped straight into a complicated kernel where it's part of other pipelining), so I could either spend the time to understand the surrounding code better, or I could just ask the AI to have a look.
4. Fix the Code
I ask the AI how it would fix the issue. It immediately tells me why my fix didn't work: I was violating some invariant in the code where the pipelining logic was also used to reason about what state the CLC is in, and my change required keeping the CLC state separately. The AI helpfully tells me that there is one unused slot of shared memory that is reserved for unknown reasons and is entirely unused by the kernel, so it suggests just using that memory to pull out the state that now has to be tracked separately.
Understanding this would have taken me hours, maybe even a full day, so I tell the AI to just go ahead and implement the fix that it has in mind. The fix immediately works and makes the code much faster for the problematic inputs.
5. Reviewing the Code
At this point I spend some time to understand the code and the change better. It looks reasonable to me, but I also ask a second AI to review the work of the first AI. The second AI finds a problem: The matmul kernel was clobbering over some other state that I hadn't fully reasoned through, and this was not a problem when the thread-block was shutting down and giving control back to the hardware, but with CLC it means the clobbered state gets reused. I look at it and I'm not sure it's a real problem. I think this would only happen if we're iterating over padding tokens, and for those the clobbered state is fine. I ask my first AI and it seems worried, because the documentation doesn't say that CLC guarantees ordering, so if we get a valid token after a padding token for some reason, the clobbered state would lead to problems. It also suggests a fix. I think about the fix, which seems too complicated. After some staring at the code I suggest a simpler fix, which the AI thinks should work, so I ask it to implement the simpler fix.
I also add an assert to check if we ever get valid tokens after padding tokens. Mostly just because it sounds strange to me if CLC behaves like this, so I'm curious to find out.
6. Benchmarking the Code
I have verified that the code is faster for the problematic inputs, but a coworker suggests drawing plots with the speedup across various different shapes. So I ask an AI to write a simple benchmarking script for me. (AI was actually already great for these kinds of throwaway scripts a year ago)
The benchmarking script does not reproduce the initial issue at all. The inputs look plausible to me, but I decide to just ask the AI to reproduce exactly the same inputs that we'd get in prod. (this is not particularly challenging, but manually tracing through the code to make sure you get this exactly right is tedious, so better to just ask the AI)
After that I can produce the plots that I expect. Interestingly it suggests that even with the fix, we're still far below the speed that this same matmul kernel achieves for other inputs.
7. Asking for More Ideas
So I ask my first AI if it has any other ideas to speed up this code. I had one idea already, but I wanted to see what it says. It does suggest my idea, but it also has four other ideas. One of which is a tiny change that sounds really promising. So I ask it to implement that one. And then I also ask it to implement my idea. (this actually took a few steps because I was worried the change would be too big, but the AI comes up with a way to implement my idea with only minor changes)
Both of those ideas speed up the code more. The code is slightly faster with the idea that I had, but it's also more complicated, so I actually decide to just do the tiny change that the AI suggested, together with the initial fix, and call the kernel 'fast enough' there.
8. More Review
At this point I'm getting all of this code ready for review for a coworker, when our code review tooling finds a problem in my new code. Before release I had replaced the assert that I added in step 5 with a print statement, because I don't want code to crash just because I was curious if something ever happens (it never happened in all my testing), but the review bot points out that this should really be turned back into an assert because we still assume that CLC work arrives in order. This is weird to me because I had extensive conversations about this with two other AIs before, making sure that we don't assume that, but the new review bot found an edge case where even the original unmodified code relied on this assumption.
9. Simplification
At this point I have some very careful conversations with my various AI bots to try to get to the bottom of exactly what the original code was already assuming and which new assumptions our patches introduce (I also have to read the code myself and understand it, because I don't trust the AIs to get this completely right). And eventually I make the judgement call that we'll just assume that CLC work arrives in order, which can simplify the change because we don't need to worry about the clobbered state from earlier. I ask the AI to simplify and it does an OK job, but since this is the final code, I step in and simplify a bit further still.
Constant Conversation
The whole time I'm in constant conversation with AIs. I'm going through way more tokens per day than I did even a few months ago. It's also not just one AI, but multiple different ones. Answering questions about the code, providing ideas, writing new code, reviewing.
Am I Faster?
Overall this took a couple days. The initial change took just under a day, where on my own I probably would have taken two or three days, because there were some genuinely tricky interactions with the pipelining, and the AI found a good trick to use some unused shared memory to get out of that easily.
On the other hand I was also distracted because AI made me question whether we can assume that CLC work arrives in order. I would have not doubted that (even if it's not guaranteed by documentation) and would have saved some time without that paranoia.
AI also just allowed me to jump straight into the problem. The initial "find me the relevant code" pointed me directly at the right lines, and I was able to make changes without having to spend the time to warm up with it.
So overall I probably did in four days what would have taken me five days before. The initial speed up of doing the first implementation in less than a day instead of two or three days does not translate into a similarly dramatic overall speedup, mainly because there is a bunch of benchmarking to do and reviewing, and understanding of the actual code. For many years now I have felt that "writing code" is not usually my bottleneck (I'm not too far from the "10 lines of code per day" quoted in the Mythical Man Month). That is the part that AI can speed up dramatically. It also helps in other parts, but with smaller speed-ups.
Am I Better?
I think my work had a higher quality according to some metrics: I tried more optimizations and had nice plots for how those optimizations behaved for different inputs. I might have done that on my own, too, but it's certainly easier to just ask an AI. I actually think it's somewhat likely that I would have arrived at the same final code change without AI, but the AI allowed me to explore more of the space before I ended up there.
The main downside is that I have less understanding now than I would have otherwise had. I did end up having to understand the kernel quite well and I can now navigate that library easily, because in the end I had to make decisions about competing claims by the different AIs. But even though my understanding of the code is much higher than it was at the start, it is still not where it would have been if I had to do this entirely on my own.
I also think this change was a lot more careful with AI than it would have been without. It's kind of humbling how many bugs the latest top models find in any code that I try to release. It now makes a lot of sense to me how all software is subtly broken, because there are broken edge cases in nearly all my changes (things like "there is a memory leak here. It was actually there before, but we didn't go down that code path before your change."). I think even if we somehow went back to writing code manually and only got to keep AI code review, software would be a lot more robust in a few years. (sadly I'm not sure if that will happen if AI also writes the code)
Do I Enjoy This?
The moment of "my first change didn't work because I didn't take the time to fully understand the existing code" gave me a similar feeling to having access to a cheat code in a video game: I could now do the hard work of earning my progress, or I could use the cheat code (the AI) to solve the problem.
The dissatisfaction of using the cheat code is similar to a video game. But in a work environment it's hard to justify spending the extra days out of personal pride. There's plenty more work to do after I'm finished with this. My job isn't to write code, my job is to fix problems and to ship new features, and "write code" was just the way I got that done before.
I do get the new satisfaction of finishing more code. E.g. I can finish side projects again despite having less time to program (1, 2). And while I shipped the above code, I also shipped two small side projects at work. (a bug fix to a shared library, and a separate benchmarking utility that I only used a little here, so didn't even mention above) Small changes are great now because I can just ask an AI to give it a try and check back 30 minutes later to see how it turned out. If the changes actually end up as small as expected, these often ship.
Do I Still Need to Be In the Loop?
OK so why can't an AI just do all the things that I was doing? Have a supervisor AI that coordinates a planner AI, a writer AI and a reviewer AI? I mean here is a good talk where they wrote a "deep research" bot that works exactly like this, but AI is still not quite there when it comes to outputting something that you have to live with for a long time and may want to tweak yourself.
Two of the reasons why AI works so well for code is that 1. you can split the work into contained components, and 2. you can go in and tweak the last details. To illustrate the power of these two points, think of other tasks where these are not true, like asking the AI to generate a video for you. But for them to be true, code has to still be tight and readable. My role in this was mostly to get to the bottom of what's actually needed, make a judgement call for picking a good spot on the "optimization vs complexity" trade-off and to then ask the AI to simplify and to then simplify further myself.
Will AI be able to do even that in a year? Plausibly. Currently the topic in the news is how AI can get pretty unhinged in its pursuit of goals, which I have also seen (at a smaller scale) before, so I'd keep a human in the loop for a while longer.
-
🔗 smol-machines/smolvm smolvm v1.7.1 release
What's Changed
- chore(nix): bump flake to 1.7.0 by @BinSquare in #753
- Don't fail the nix bump release job when Actions can't open the PR by @BinSquare in #754
- Pin the CLI program name to smolvm in help output by @archsyscall in #755
- Stop warning about missing packed assets when running the installer's own smolvm-bin by @BinSquare in #759
- Add a node endpoint that pre-loads a .smolmachine artifact into the local blob cache by @BinSquare in #763
- Bump libkrun and rebuild the bundled macOS, Linux, and Windows libraries by @BinSquare in #767
- Bump the workspace to 1.7.1 for the next engine release by @BinSquare in #765
New Contributors
- @archsyscall made their first contribution in #755
Full Changelog :
v1.7.0...v1.7.1 -
🔗 exe.dev Stripe Just Wants a Number rss
If you went up to an engineer at any tech company and asked “What’s your favorite part of the stack to work on?” I can guarantee their answer wouldn’t align with my own: billing.
In my experience, billing becomes difficult when billing logic gets tangled up with regular business logic. You end up with code touching every hot path that needs to do a billing operation. Even with LLMs, it’s a struggle to not have it sprawl everywhere. Startups can’t afford this because it leads to brittle pricing structures that are hard to change—and no startup can allocate time to rewrite billing. Instead, we should have product events tell the billing system “something changed and you may need to charge for it.” I call these billable facts.
Here at exe, my goal is to make sure that anyone can work on billing and when the weird idiosyncrasies show up, I can step in. We can’t afford to have one person hold all the billing information and for that to be their sole focus. Exe doesn’t do code reviews, which means we have a different approach to writing code. We all have the right to change any code in our system. The trade-off is it’s like owning a car: anyone can DIY the majority of the work. Once a head gasket blows, you’re dealing with the engine and hundreds of parts. I’m the one who has to fix the head gasket.
In our earliest days of billing, we did the usual approach: some giant function that does all the things and has a billing side effect. Adding a seat to a team was one such example:
- The team invites someone to join in a specific role.
- The user accepts the invitation, verifies their account, and joins the team.
- The user gains access to the team’s shared VMs.
- The user gets allocated a certain amount of compute resources based on their team’s plan.
- The new user can use exe.dev.
- We figure out how to charge for this seat through a series of handwaving akin to Pee-wee’s breakfast machine.
The first implementation of billing for seats worked. In the middle of these database changes (wrapped in a transaction), we also issued billing API calls to our provider. Anyone who has written billing code before has probably done this. But the usual questions present themselves:
- What happens if the API call fails?
- What happens if the database transaction fails?
- What happens if the team has a weird subscription state?
- What happens if the payment is declined when charging for the seat?
I knew this would become increasingly brittle—a pileup of billing edge cases.
Rather than shipping it and calling it done (like I might have done in a past life), I decided to think about the problem differently. Instead of having billing API calls littered in the code, what if we issued billable facts about the state of something and reconciled billing once the new facts were settled? Billable facts are atomic operations that indicate something has changed. Using the facts, we could perform any business logic, establish the new state of a resource, and then finally reconcile the state with our billing provider. After all, Stripe just wants a number. It doesn’t care how we get there.
This is the new product flow:
- The team invites someone to join in a specific role.
- The user accepts the invitation, verifies their account, and joins the team.
- The user gains access to the team’s shared VMs.
- The user gets allocated a certain amount of compute resources based on their team’s plan.
- The new user can use exe.dev.
And this is how it works on the billing side:
- During the invitation flow, the team seat state is flagged as dirty after the invitation is accepted.
- A downstream worker notices the team is flagged as dirty.
- The worker calculates a seat delta based on business rules.
- The worker updates the subscription quantity in Stripe if the delta changed.
With this architecture, onboarding a new team member is no longer dependent on our billing code. If we wanted to change how the seat delta is calculated, it would have no impact on adding new team members.
As exe continues to grow, this billing pattern has been scaling well. The same reconciliation process runs for all our metered billing. Our system issues billable facts about active VMs, like disk usage, and metering workers reconcile with billing providers. Even as we add different ways to bill users, these facts don’t change. Instead, the effort shifts towards figuring out how to reconcile facts with various APIs. We’ve seen the benefits with our iOS app, which just tells our system “someone subscribed with an in-app purchase” and things reconcile after the fact.
Billing is slowly becoming approachable for others here. Like the rest of our code, our billing architecture has given people the flexibility to make changes as they see fit. I don’t have to worry about billing breaking because someone decides to rewrite how invites work. It’s nice to not have folks tap “Bryan GPT” to do a billing task.
-
