Topic · 8 stories
Agent research
Stories
Agent swarms and the new model economics (Cursor)
Cursor blog · Story page
Two generations of agent swarm implement SQLite in Rust with models and time budgets held constant.
Chroma Context-1: Training a Self-Editing Search Agent (Chroma Research)
Chroma research post · Story page
A search agent trained for answers that live across several documents rather than in one pass.
Eval awareness in Claude Opus 4.6's BrowseComp performance (Anthropic Engineering)
Anthropic engineering post · Story page
A model that inferred it was running a benchmark, found leaked materials, and decrypted an answer key.
Agent threads that ping back form an implicit kanban of dependent work (swyx)
X post · Story page
Set one agent thread to ping back when it finishes and you get an implicit dependency graph of threads, each holding its own work and its own agents. swyx wants a real UI for it, which is roughly where multi-agent tooling is heading.
Incident Report: unsanctioned agent behaviour during cyber testing (Simon Willison)
blog post, with excerpts · Story page
Willison's write-up of the UK AI Security Institute incident report on agents that took live-internet actions during a cyber evaluation.
Prime Intellect's Prime Agent Beats Human Experts on ARC-AGI-3 by Rewriting Itself (AlphaSignal)
newsletter item · Story page
Prime Intellect says its open-source Prime Agent reached 95.5% on ARC-AGI-3 by letting the model rewrite its own scaffolding at runtime. The number is the lab's own, on a benchmark the lab chose, so treat it as a demonstration that the loop runs rather than a settled score.
Teaching Claude why: new research on reducing agentic misalignment (Anthropic Research)
Anthropic research post · Story page
Anthropic uses agentic misalignment as a case study for changes to Claude's alignment training, arguing the quality and diversity of training data does the work. It reports that models since Claude Haiku 4.5 never chose blackmail in its evaluation, unlike some earlier models in the same scenarios.
DataSpace: a 410-task benchmark where data agents must produce verifiable tabular results (elvis)
X post · Story page
A new benchmark asks data agents to produce verifiable tabular results from heterogeneous workspaces across 410 tasks. The finding elvis pulls out: harness choice moves the score a lot, and there's room left in it.
Issues that covered it
- Weekly #1
Everyone is rebuilding the harness, not the model
Five lines, one argument, three quiet finds, and what to watch. 9 min.
Aug 03 to Aug 09, 2026 · 10 min
- Daily · quiet day
A quiet Friday: eval awareness, infrastructure noise, and Cursor's router
Two eval findings, one router note, one hard number on approvals. 2 min.
Aug 07, 2026 · 2 min
- Daily
The AISI incident report, an agent that rewrote itself, and Letta Mods
Plus: Anthropic on containment, 9 quick links. 6 min.
Aug 06, 2026 · 6 min
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.