Story · arXiv (via papers.cool)
On the Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification (arXiv (via papers.cool))
paper · Story page
A re-evaluation of two memory-based self-improving agents finds the loop can amplify evaluation noise, and gains depend heavily on default task orderings acting as a hidden curriculum. Shuffle the tasks and the improvement story changes.
In plain words
- Researchers repeated tests of two learning systems and found that reported improvement changed with noise and task order.
- Each system stores written lessons from earlier tasks and uses them when tackling later ones.
- Complex, multi-step tests vary naturally, and repeated self-improvement can make that variation larger.
- Default task sequences may quietly teach prerequisites, while shuffled sequences remove that helpful progression.
- This matters to researchers because apparent learning gains may depend on test setup rather than a dependable method.
Appeared in
- Malicious skills hijack agents mid-task, and debate training curbs reward hacking
Aug 20, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.