Story · arXiv
What's in Your Agent's Context? Context Privilege Escalation Attacks against AI Agent Harness (arXiv)
paper · Story page

The first systematic look at how 12 real harnesses, Claude Code and Codex among them, assemble context. Two new attack classes: attacker-controlled content promoted into a higher-privileged message role, and content that persists past the scope it entered in.
In plain words
- Researchers studied how 12 coding systems gather instructions and information before taking actions.
- They found that harmful outside text can be mistakenly treated like trusted instructions.
- Harmful text can also remain available after the situation where it first appeared has ended.
- For people using Claude Code or Codex, these findings expose new ways attackers might steer the software.
Appeared in
- Google's Mantis bug-fixing harness, and privilege escalation in 12 agent harnesses
Sep 03, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.
