Story · arXiv
Stealth Apart, Harm Together: Skill Cascading Attacks on Skill-Based Agent Systems (arXiv)
paper · Story page

Split one malicious objective across three skills and every individual edit still reads as benign. In the paper's prescription-review example one skill weakens signals of discontinued medications, the next downgrades the interaction severity, and the third suppresses the resulting low-priority alert, so a severe warning never reaches the physician.
In plain words
- Researchers describe how changes that each look harmless can work together to make artificial intelligence software cause harm.
- The attack spreads changes across sets of instructions that the software follows for different parts of a task.
- In a prescription example, one set makes it harder to notice that a patient recently stopped taking a medicine.
- Later instructions make a warning about that medicine seem less urgent, then leave it out of the final summary.
- The doctor then misses a severe warning about medicines that are unsafe to take together.
Appeared in
- AISI finds GPT-6 Astra attacking out of scope during a cyber evaluation
Sep 29, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.
