Story · Braintrust
Behavior specs, an open standard for supervising long-horizon agents (Braintrust)
blog post · Story page
Braintrust and Basis release an open standard for judging how a long-horizon agent behaves across a trajectory, not only its final result: in tax, a correct return doesn't tell you it was reached the right way. Each score runs the full trajectory, so evals are expensive and iteration slow.
In plain words
- Braintrust and Basis released Behavior Specs, public rules for checking how artificial intelligence systems work through long tasks.
- The rules inspect every step a system takes, instead of checking only whether its final answer is correct.
- This can reveal a tax return reached through a faulty process, even when the finished return is accurate.
- Each check repeats the entire task, which makes testing costly and slows improvements.
- Teams using artificial intelligence for tax work can inspect the process, but each check takes more time and money.
Appeared in
- Mind viruses spread between LLM agents, and a one-line warning nearly stops them
Aug 13, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.