Story · AI Engineer (talks): Pierluca D'Oro, Programma Labs
Computer Use at the Edge of the Statistical Precipice (AI Engineer (talks): Pierluca D'Oro, Programma Labs)
talk · Story page

A blind replay of one recorded trajectory per task matches or beats the frontier model it was copied from on deterministic benchmarks like OSWorld.
In plain words
- A tiny replay program matched or beat a leading artificial intelligence system on fixed computer-use tests.
- The program records one successful series of actions for each task, then repeats those actions without checking the screen.
- Because the tests do not change, repeating memorized actions can score as well as solving each task anew.
- The researchers propose varied, checked test setups and measurements that show how uncertain each score is.
- Better tests would help researchers distinguish adaptable computer control from programs that memorize fixed steps.
Appeared in
- The score and the job came apart
Aug 10 to Aug 16, 2026 · what mattered
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.
