Story · arXiv (via papers.cool)
Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills (arXiv (via papers.cool))
paper · Story page
ACES answers what a static skill scan can't: does the package help a live agent? It runs paired trials with and without each skill under the same model, sandbox, tasks, and scorer, then reports the measured lift, evaluated on 145 real skills.
In plain words
- Researchers introduced Agentic Continuous Evaluation of Skills, a system testing whether reusable instructions and tools help artificial intelligence complete workplace tasks.
- It compares the same task twice, once with the package of instructions and tools and once without it.
- Both attempts use the same artificial intelligence, isolated workspace, tasks, and grading rules, making the package the main difference.
- Companies can identify packages that improve workplace results instead of trusting descriptions or checks of file structure.
Appeared in
- Poisoned agent memory beats screening, and prompt rules keep failing as boundaries
Aug 25, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.