Story · arXiv
Hearsay: Can an Auditor Trust the Record a Deployed Agent Harness Writes? (arXiv)
paper · Story page
Across sixteen deployed frameworks, none writes a complete run record a reader can check without trusting whatever wrote it. Examiners named the right fault in 74 to 91% of 140 runs, but under one citation in ten about intermediate events landed on anything the harness didn't write.
In plain words
- Researchers found that investigators could not fully check what artificial intelligence assistants had done using the records provided.
- The software controlling each assistant writes its own activity record, so investigators must trust that software when reading it.
- Readers often identified problems correctly, but lacked independent evidence for much of what happened along the way.
- The researchers propose independently kept records so investigators can check disputed actions without relying on the software's own account.
Appeared in
- Agent context compression can cut tokens by two thirds and still run slower
Sep 30, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.