Story · arXiv

The Off-Support Barrier: Why Semantic Safety Constraints Are Not Learning-Problem Invariants, and What Follows for Prior Design, Containment, and Verification (arXiv)

paper · Story page

A spotlight lights a circle on a dark floor holding a small figure; a sign with a crossed-out padlock stands outside the lit area, and a fence surrounds everything.

A theory paper argues that constraints like 'the agent does not escape its sandbox' aren't measurable from the model and training distribution alone, so training can't guarantee them. Hard invariants belong in the harness and formal verification; the model carries soft dispositions.

In plain words

  • The paper argues that training alone cannot guarantee rules such as keeping an action-taking artificial intelligence system inside a restricted environment.
  • Training covers expected situations, while safety rules may concern unfamiliar situations outside that experience.
  • Changing training preferences or adding soft penalties provides little control over these hard rules, according to the paper.
  • Separate control software can enforce fixed limits, while mathematical checking can confirm specific rules in defined situations.
  • Builders therefore need safety controls outside the trained system when failures such as escape must be prevented.

Appeared in

Subscribe

Get the brief in your inbox

Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.

  • Weekdays at 8:45am IST, one lead story and 6 to 9 items.
  • Sundays, an argued synthesis rather than a recap.
  • One click to leave, and quiet days say so in the subject line.
How often

Weekdays 8:45am IST + Sundays. Unsubscribe in one click.

You're asking for The Agentic Brief by email at the cadence you picked. You can unsubscribe in one click from any issue, and your address is never sold or shared.