Story · Semantic Scholar
Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents (Semantic Scholar)
paper · Story page
A differential method pins a failed or costlier run on the skill it loaded by comparing it with a no-skill or matched-skill run of the same task. On SkillsBench and SWE-Skills-Bench that yields 307 skill-induced failures, 125 functional and 182 efficiency.
In plain words
- A study found 307 cases where reusable instructions caused artificial intelligence to fail tasks or work less efficiently.
- Researchers compared each instructed run with the same task completed without those instructions or with a similar set.
- These side-by-side comparisons show what changed and whether the instructions caused failure or extra cost.
- SkillTriage organizes the differences into consistent reports supported by evidence from both runs.
- Teams adding reusable instructions can use this method to check whether they improve results or quietly increase costs.
Appeared in
- Deno's Claw Patrol treats agents as untrusted software, and full history beats compaction
Aug 18, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.