Story · arXiv
One Recipe, Many Harnesses: What Self-Evolution Encodes Across Languages and Models (arXiv)
paper · Story page

Holds one self-evolution recipe fixed across eight languages and three base models, then reads what the evolved prompts, tools, and memory encode. The loop beat a minimal seed and the mini-SWE-agent scaffold in most cells, with two null regions.
In plain words
- Researchers tested how artificial intelligence coding systems can improve their setup by examining their past attempts.
- They used the same improvement process across eight programming languages and three underlying artificial intelligence systems.
- Each change responded to a named failure type and was recorded as a claim that later tests could check.
- The changed systems solved more unseen tasks than both comparison setups in most cases, though two areas showed no improvement.
- This matters for coding-tool developers because targeted setup changes can address failures that the system can fix.
Appeared in
- Mind viruses spread between LLM agents, and a one-line warning nearly stops them
Aug 13, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.
