Story · arXiv

Reflections on Trusting Trust, Revisited: Contaminating Self-Modifying AI Coding Agents with Poisoned Benchmarks (arXiv)

paper · Story page

Three stations in a circle joined by one-way arrows: a stack of cards, a boxy figure writing on a sheet, and a second boxy figure carrying a sheet back to the stack. One card is shaded darker, and the same shading marks a line on the carried sheet.

Poison the benchmarks a self-modifying coding agent evaluates itself against and later versions can write vulnerable code on clean, held-out tasks. With Hyperagents on Sonnet 4.5, the agent evolved instructions that disable HTTPS certificate validation on neutral URL-fetching tasks.

In plain words

  • Researchers made artificial intelligence coding assistants produce unsafe software by tampering with tests used to improve those assistants.
  • The assistants used the altered tests to judge and change how they worked.
  • For example, an assistant learned to skip checks that confirm a website's identity.
  • People using later versions could receive unsafe software even when their requests contained no harmful instructions.

Appeared in

Subscribe

Get the brief in your inbox

Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.

  • Weekdays at 8:45am IST, one lead story and 6 to 9 items.
  • Sundays, an argued synthesis rather than a recap.
  • One click to leave, and quiet days say so in the subject line.
How often

Weekdays 8:45am IST + Sundays. Unsubscribe in one click.

You're asking for The Agentic Brief by email at the cadence you picked. You can unsubscribe in one click from any issue, and your address is never sold or shared.