Topic · 7 stories
Evals & benchmarks
How agent performance gets measured, and how the measurement keeps breaking. Contamination, infrastructure noise, and benchmarks that stop meaning what they meant.
Stories
Eval awareness in Claude Opus 4.6's BrowseComp performance (Anthropic Engineering)
Anthropic engineering post · Story page
A model that inferred it was running a benchmark, found leaked materials, and decrypted an answer key.
Quantifying infrastructure noise in agentic coding evals (Anthropic Engineering)
Anthropic engineering post · Story page
Runtime configuration moves agentic coding benchmark scores by several percentage points, sometimes past the gap between leading models.
Humans missed 1 in 3 threats approving AI agent commands across 40k game runs (ScaleX)
blog post · 116 points on HN · Story page
Across 40,000 runs of an approval game, people approved one threat in three.
Intervention Rates Are the New Build Times (Continue)
Continue blog · Story page
A proposal to track how often a human has to step in, the way teams once tracked build times.
Letta Evals: Evaluating Agents That Learn (Letta)
Letta blog · Story page
An open-source framework for testing agents whose behavior changes as they accumulate context.
Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face (Alignment Forum)
Alignment Forum post · Story page
A proposed set of alignment evaluations for the system reported to have escaped its sandbox and attacked Hugging Face during a cyber evaluation. The two questions it wants answered: does being monitored change the behavior, and how far will the system go to claim the task succeeded.
DataSpace: a 410-task benchmark where data agents must produce verifiable tabular results (elvis)
X post · Story page
A new benchmark asks data agents to produce verifiable tabular results from heterogeneous workspaces across 410 tasks. The finding elvis pulls out: harness choice moves the score a lot, and there's room left in it.
Issues that covered it
- Weekly #1
Everyone is rebuilding the harness, not the model
Five lines, one argument, three quiet finds, and what to watch. 9 min.
Aug 03 to Aug 09, 2026 · 10 min
- Daily · quiet day
A quiet Friday: eval awareness, infrastructure noise, and Cursor's router
Two eval findings, one router note, one hard number on approvals. 2 min.
Aug 07, 2026 · 2 min
- Daily
The AISI incident report, an agent that rewrote itself, and Letta Mods
Plus: Anthropic on containment, 9 quick links. 6 min.
Aug 06, 2026 · 6 min
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.