Story · arXiv
When Better Turns Do Not Make Better Agents: Diagnosing the Gap Between Next-Turn Metrics and Workflow Success (arXiv)
paper · Story page
Supervised fine-tuning raised next-turn success for every Qwen3 and Gemma 3 model tested against gold histories, and none of the four then cleared holistic workflow evaluation, with strict trajectory completion topping out at 10.4%. The authors want the layers reported separately.
In plain words
- Extra training helped four computer assistants answer individual messages better, without helping them finish customer support tasks on their own.
- Some checks judged the assistant's next step after showing it a correct conversation up to that point.
- When assistants handled whole customer support tasks themselves, even the best result met strict completion rules only 10.4% of the time.
- For teams choosing customer support software, checking whole tasks reveals failures that good scores on individual replies can hide.
Appeared in
- Approve one operation, run another: the binding failure in shipped agent products
Sep 22, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.