Topic · 8 stories
Harnesses & tools
The code around the model: tool access, context handling, subagent orchestration, and the harness-level features that increasingly decide what an agent can finish.
Stories
What Is a Skill Worth? Structure-Aware Shapley Valuation of Agent Skills (arXiv)
arXiv preprint · Story page
SkillSV puts a value on the pieces inside an agent skill, separating what the content is worth from what the context costs.
Microsoft's Orchard Beats Proprietary AI Agents at 10x Lower Cost (AlphaSignal)
newsletter item · Story page
A Kubernetes-native framework for training coding, browser, and productivity agents with small models.
Building a C compiler with a team of parallel Claudes (Anthropic Engineering)
Anthropic engineering post · Story page
Sixteen parallel agents built a Rust C compiler able to compile the Linux kernel, with little active supervision.
Agent threads that ping back form an implicit kanban of dependent work (swyx)
X post · Story page
Set one agent thread to ping back when it finishes and you get an implicit dependency graph of threads, each holding its own work and its own agents. swyx wants a real UI for it, which is roughly where multi-agent tooling is heading.
How Cursor Router chooses the right model for the task (Cursor)
Cursor blog · Story page
Cursor's router is trained on performance against real developer work rather than benchmark scores, and picks a model from the current turn, recent conversation state, task category, and tool calls. Model routing is quietly becoming a harness responsibility rather than a settings toggle.
Prime Intellect's Prime Agent Beats Human Experts on ARC-AGI-3 by Rewriting Itself (AlphaSignal)
newsletter item · Story page
Prime Intellect says its open-source Prime Agent reached 95.5% on ARC-AGI-3 by letting the model rewrite its own scaffolding at runtime. The number is the lab's own, on a benchmark the lab chose, so treat it as a demonstration that the loop runs rather than a settled score.
Introducing Mods: Enabling Agents to Self-Improve through Harness-Level Adaptation (Letta)
Letta blog · Story page
Mods lets an agent extend and revise the Letta Code harness itself, not only its prompts, memory, or skills. It treats harness behavior as learnable state, building on Letta's versioned context store and self-editing memory tools.
Harness design for long-running application development (Anthropic Engineering)
Anthropic engineering post · Story page
Generator-evaluator loops, task decomposition, and structured context handoffs, ending in a planner, generator, and evaluator setup.
Issues that covered it
- Weekly #1
Everyone is rebuilding the harness, not the model
Five lines, one argument, three quiet finds, and what to watch. 9 min.
Aug 03 to Aug 09, 2026 · 10 min
- Daily · quiet day
A quiet Friday: eval awareness, infrastructure noise, and Cursor's router
Two eval findings, one router note, one hard number on approvals. 2 min.
Aug 07, 2026 · 2 min
- Daily
The AISI incident report, an agent that rewrote itself, and Letta Mods
Plus: Anthropic on containment, 9 quick links. 6 min.
Aug 06, 2026 · 6 min
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.