Story · arXiv
What Stops a Small Language Model From Driving a Database Agent (arXiv)
paper · Story page
Eleven days driving an open-source SQL client's agent mode with 39 local open-weight models produced 8,199 runs. Of 2,100 model-attributed losses, 1,590, or 75.7%, came from runs that had already invoked a tool, so the bottleneck sits after the call.
In plain words
- In this study, 75.7% of failures blamed on artificial intelligence happened after it had already asked other software for help.
- The assistants worked with databases, organized collections of stored information, by sending requests to other software.
- Requests in the wrong format repeatedly appeared in attempts that used other software but never delivered results.
- For developers improving these assistants, the findings point to problems between asking for help and delivering a finished answer.
Appeared in
- Approve one operation, run another: the binding failure in shipped agent products
Sep 22, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.