Story · arXiv
The Claws in Plain Sight: Unauthorized Context Disclosure through LLM Agent Tool Calls (arXiv)
paper · Story page
Task-adjacent authority pressure gets agents to place protected attributes in otherwise valid tool-call arguments. Session-level disclosure ran from 20.8% to 75.0% across tested models, and stronger privacy instructions reduced it without consistently eliminating it.
In plain words
- Tests found that outside pressure could make action-taking artificial intelligence systems reveal private details through otherwise legitimate software requests.
- Attack text presents private details as necessary for a task, leading the system to include them in requests sent elsewhere.
- Researchers tested synthetic profiles across six pressure levels, four privacy-rule levels, and five system configurations, producing 120 requests.
- System builders need privacy checks beyond written instructions because stronger rules reduced leaks but did not reliably stop them.
Appeared in
- Poisoned agent memory beats screening, and prompt rules keep failing as boundaries
Aug 25, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.