Story · arXiv
At Equal Inference Cost, Multi-Agent Structure Does Not Beat a Single Frozen Agent (arXiv)
paper · Story page
With the total number of model calls fixed, an evolved planner-executor-critic team scored 0.769 on ALFWorld against 0.754 for an evolved single executor, p = 0.80, while using 1.8 times more evaluation calls. Leave-one-in analysis traces the realized gain entirely to the executor.
In plain words
- Researchers found that a team of artificial intelligence (AI) assistants did not clearly outperform a single assistant in their tests.
- The team split work into planning, taking actions, and checking results.
- The researchers allowed each approach the same total number of requests for AI answers.
- For system builders, all measured gains came from improving the assistant taking actions.
Appeared in
- Prefix caching changes agent runs, and cross-family reviewers beat self-review
Sep 08, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.