Story · arXiv
The Off-Support Barrier: Why Semantic Safety Constraints Are Not Learning-Problem Invariants, and What Follows for Prior Design, Containment, and Verification (arXiv)
paper · Story page

A theory paper argues that constraints like 'the agent does not escape its sandbox' aren't measurable from the model and training distribution alone, so training can't guarantee them. Hard invariants belong in the harness and formal verification; the model carries soft dispositions.
In plain words
- The paper argues that training alone cannot guarantee rules such as keeping an action-taking artificial intelligence system inside a restricted environment.
- Training covers expected situations, while safety rules may concern unfamiliar situations outside that experience.
- Changing training preferences or adding soft penalties provides little control over these hard rules, according to the paper.
- Separate control software can enforce fixed limits, while mathematical checking can confirm specific rules in defined situations.
- Builders therefore need safety controls outside the trained system when failures such as escape must be prevented.
Appeared in
- Agent leaderboards rank specialization, and the harness thesis reaches robots
Aug 14, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.
