Story · METR
Implementing and Evaluating a Basic Per-Action Monitor for Safer Evals (METR)
METR put an LLM judge in front of every agent action in its own evaluations, holding anything above the risk threshold for human review and halting the eval until someone looks. It's candid about where the evidence falls short.
In plain words
- Researchers added a safety check before every action taken by artificial intelligence assistants during their tests.
- Another artificial intelligence system reviews each planned action before it happens.
- Actions judged too risky pause the test until a person reviews them.
- The team describes missing evidence about how well the checks work.
- Researchers aim to reduce the risk of real harm without overwhelming the people reviewing potentially risky actions.
Appeared in
- AISI finds GPT-6 Astra attacking out of scope during a cyber evaluation
Sep 29, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.