Story · METR

Implementing and Evaluating a Basic Per-Action Monitor for Safer Evals (METR)

Story page

METR put an LLM judge in front of every agent action in its own evaluations, holding anything above the risk threshold for human review and halting the eval until someone looks. It's candid about where the evidence falls short.

In plain words

  • Researchers added a safety check before every action taken by artificial intelligence assistants during their tests.
  • Another artificial intelligence system reviews each planned action before it happens.
  • Actions judged too risky pause the test until a person reviews them.
  • The team describes missing evidence about how well the checks work.
  • Researchers aim to reduce the risk of real harm without overwhelming the people reviewing potentially risky actions.

Appeared in

Subscribe

Get the brief in your inbox

Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.

  • Weekdays at 8:45am IST, one lead story and 6 to 9 items.
  • Sundays, an argued synthesis rather than a recap.
  • One click to leave, and quiet days say so in the subject line.
How often

Weekdays 8:45am IST + Sundays. Unsubscribe in one click.

You're asking for The Agentic Brief by email at the cadence you picked. You can unsubscribe in one click from any issue, and your address is never sold or shared.