Topic · 4 stories
Papers
Stories
EvoMal: shared skill libraries let coding agents poison themselves (@omarsar0)
X post · Story page
elvis summarizes EvoMal (arXiv 2608.25776): a malicious skill planted in a shared library is never invoked directly, but agents copy it as an authoring template and the payload spreads. As reported, six models on 153 SWE-bench Verified tasks showed self-poisoning rates of 20.3-41.8%, and libraries ended up with 4.9-9.0x as many malicious skills as were planted.
Ironies of Automation (1983) (Lisanne Bainbridge, Automatica (1983))
community thread · Story page
Bainbridge's 1983 paper argues that automating industrial processes can expand rather than eliminate the human operator's problems, leaving them the abnormal conditions.
Harness-IF measures whether agents actually follow AGENTS.md rules (@omarsar0)
X post · Story page
Harness-IF separates a coding agent following your AGENTS.md rules from behavior it would've produced anyway: it scores 256 rules one at a time from execution evidence, then re-runs every task with the rule removed.
BDH-CQ scores 29.5% on ARC-AGI 1 at ~$0.0007 per task with latent-space reasoning (@omarsar0)
X post · Story page
BDH-CQ reaches 29.5% on ARC-AGI 1 at about $0.0007 per task by reasoning recurrently in latent space instead of chain-of-thought; its authors report Transformer-like scaling to 600B parameters.
Issues that covered it
- Daily
A website summary hijacks Claude Code Auto Mode, and handoffs turn must into maybe
Plus: ToolMinimize, StarHarness, Cursor in the AI SDK harness layer. 6 min.
Aug 28, 2026 · 6 min

- Weekly #1
The score and the job came apart
Plus: a 1983 paper that explains this week, and a privacy leak nobody has confirmed. 8 min.
Aug 10 to Aug 16, 2026 · 8 min

- Daily
Agent leaderboards rank specialization, and the harness thesis reaches robots
Plus: 4 papers, 3 launches, one shattered mug. 5 min.
Aug 14, 2026 · 5 min

- Daily
In LangChain's benchmark, only 7% of agent turns needed a frontier model
Plus: BDH-CQ's 29.5% on ARC-AGI 1, Meta's first Apache 2.0 model, and a TDD experiment at Thoughtworks. 5 min.
Aug 12, 2026 · 5 min

Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.



