Story · arXiv
Recursive self-improvement of AI research agents (arXiv)
paper · Story page
AIDE^2 proposes changes to its own research-agent code, benchmarks the modified versions of itself on AI R&D tasks, and keeps whatever wins on hidden evaluations. An autonomous eight-day run found seven successive improvements, and the gains carried to four held-out benchmarks.
In plain words
- Researchers built an artificial intelligence research assistant that improves itself by rewriting its own software.
- It tries changes on research tasks and keeps the versions that perform best in hidden tests.
- Each improved version becomes the starting point for the next round of changes.
- The improvements carried over to four separate test sets, suggesting the assistant could help researchers with tasks beyond those used during development.
Appeared in
- Two MemOS packages shipped credential stealers into the agent memory layer
Sep 24, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.