Story · arXiv
FIRE: Failure-Informed Runtime Engineering for Reliable Language-Model Agents (arXiv)
paper · Story page

Runtime policies are natural-language instructions and action denials the harness applies at the states that preceded observed failures, leaving weights and the user prompt untouched. Across the 87-task Terminal-Bench 2.1 suite, repeated success for GPT-5.6 Sol moves from 64.4% to 73.6%.
In plain words
- Researchers made an artificial intelligence assistant more consistent at finishing tasks by giving it rules based on earlier mistakes.
- When the assistant reaches a situation linked to past failures, the software gives extra instructions or blocks particular actions.
- The approach needs no retraining of the assistant or changes to the user's request.
- In tests, one assistant completed both tries on 73.6% of tasks, up from 64.4%, suggesting more dependable help for users.
Appeared in
- Two MemOS packages shipped credential stealers into the agent memory layer
Sep 24, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.
