Story · Semantic Scholar
One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL (Semantic Scholar)
paper · Story page

Train a policy against one LLM user simulator and it overfits that simulator's dominant mode, then transfers poorly to other simulators and real users. Verbalized Sampling lifts held-out success by up to 9%; co-training against a population of simulators reaches 14%.
In plain words
- A study found that training with one simulated user made artificial intelligence perform poorly with other simulators and real people.
- The artificial intelligence learned narrow tactics that worked mainly against the simulated user's most common behavior.
- Verbalized Sampling makes the simulated user produce a wider range of responses instead of its usual pattern.
- Co-Training practices against several changing simulators, preventing the system from adapting to only one.
- Teams training human-facing artificial intelligence need varied simulated users to prepare behavior that works beyond one simulator.
Appeared in
- Deno's Claw Patrol treats agents as untrusted software, and full history beats compaction
Aug 18, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.
