Story · arXiv
Reflections on Trusting Trust, Revisited: Contaminating Self-Modifying AI Coding Agents with Poisoned Benchmarks (arXiv)
paper · Story page

Poison the benchmarks a self-modifying coding agent evaluates itself against and later versions can write vulnerable code on clean, held-out tasks. With Hyperagents on Sonnet 4.5, the agent evolved instructions that disable HTTPS certificate validation on neutral URL-fetching tasks.
In plain words
- Researchers made artificial intelligence coding assistants produce unsafe software by tampering with tests used to improve those assistants.
- The assistants used the altered tests to judge and change how they worked.
- For example, an assistant learned to skip checks that confirm a website's identity.
- People using later versions could receive unsafe software even when their requests contained no harmful instructions.
Appeared in
- Split an MCP injection across two channels and resistant models leak at 100%
Sep 18, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.
