Archive · 24 issues
The archive
Every issue, nothing gated. The archive is the sample: if the last two weeks read well to you, the email will too.
- Daily
Split an MCP injection across two channels and resistant models leak at 100%
Plus: reward hacking at 57.2% of rollouts, and Exa searches the web by date. 5 min.
Sep 18, 2026 · 5 min

- Daily
Every audited chat tokenizer lets prompt text forge control tokens
Plus: 91% of audited vibe-coded deployments shipped a hole. 5 min.
Sep 17, 2026 · 5 min

- Daily
A shell beats a typed tool catalog, and agents barely report their work
Plus: an MCP registry census, four AI Engineer talks, one very late transcription. 5 min.
Sep 16, 2026 · 5 min

- Daily
An agent deleted an AML control, and benchmark scaffolds do the model's work
Plus: OpenAI's 10,000-agent proof, a 297-iteration model loop, MCP merges Skills. 6 min.
Sep 14, 2026 · 6 min

- Daily
Per-phase model routing cuts agent cost, RASER on Slurm, Amp steers mid-run
Plus: 2 papers, a context-mode split, and OpenAI's Navier-Stokes claim. 5 min.
Sep 09, 2026 · 5 min

- Daily
Prefix caching changes agent runs, and cross-family reviewers beat self-review
Plus: a harness that moves solve rates 4x, and 883 commits to nowhere. 6 min.
Sep 08, 2026 · 6 min

- Daily
GPT-6 Astra's score hinges on its harness, and agents rot geometrically with each step
Plus: Cursor agents in Vercel microVMs, Composio's six missing primitives, ACLE-MCP. 5 min.
Sep 04, 2026 · 5 min

- Daily
Google's Mantis bug-fixing harness, and privilege escalation in 12 agent harnesses
Plus: Cline's 11M-user rollout, ContextPipe, and Claude driving your desktop. 5 min.
Sep 03, 2026 · 5 min

- Daily
Anthropic's deliberately misaligned model, Fable 5.1, and a fix for reward hacking
Plus: escalation channels cut reward hacking, memory that rots, Copilot approvals. 6 min.
Sep 02, 2026 · 6 min

- Daily
Realistic prompts drop coding-agent scores, and tool filtering beats prompt rules
Plus: 7 papers, 2 releases, one empty-handed mugger. 5 min.
Sep 01, 2026 · 5 min

- Daily
Maersk's 100,000 corrections, preference-trap evals, and Tencent's Hy4 Preview
Plus: guardrails at Navan, a one-way fiber, and the benchmark card that names its scaffold. 5 min.
Aug 31, 2026 · 5 min

- Daily
A website summary hijacks Claude Code Auto Mode, and handoffs turn must into maybe
Plus: ToolMinimize, StarHarness, Cursor in the AI SDK harness layer. 6 min.
Aug 28, 2026 · 6 min

- Daily
Compaction erases agent safety rules, and CLAUDE.md prose isn't a control
Plus: Vercel's Run SDK, Scroll's executable context, and why the brainstorm converged. 5 min.
Aug 26, 2026 · 5 min

- Daily
Poisoned agent memory beats screening, and prompt rules keep failing as boundaries
Plus: bash-sed-cat at AIDAChip, a breach in Thailand, 8 quick links. 5 min.
Aug 25, 2026 · 5 min

- Daily
v0 keeps OAuth tokens out of generated code, plus SkillGate and Temporal's harness
Plus: 4 papers, 2 talks, and a meme about durable execution. 5 min.
Aug 21, 2026 · 5 min

- Daily
Malicious skills hijack agents mid-task, and debate training curbs reward hacking
Plus: an unverified OpenAI sandbox story and a fail-closed runtime. 5 min.
Aug 20, 2026 · 5 min

- Daily
Aborted agent branches live on in the KV cache, plus StateM's harness scaling
Plus: coherence debt, the recall trap, Cline's open-weight evals. 5 min.
Aug 19, 2026 · 5 min

- Daily
Deno's Claw Patrol treats agents as untrusted software, and full history beats compaction
Plus: 6 papers, 2 talks, a 56,000-line Fortran migration. 5 min.
Aug 18, 2026 · 5 min

- Weekly #1
The score and the job came apart
Plus: a 1983 paper that explains this week, and a privacy leak nobody has confirmed. 8 min.
Aug 10 to Aug 16, 2026 · 8 min

- Daily
Agent leaderboards rank specialization, and the harness thesis reaches robots
Plus: 4 papers, 3 launches, one shattered mug. 5 min.
Aug 14, 2026 · 5 min

Browse by topic



















