Story · AI Engineer (talks): Pierluca D'Oro, Programma Labs

Computer Use at the Edge of the Statistical Precipice (AI Engineer (talks): Pierluca D'Oro, Programma Labs)

talk · Story page

A tape player feeds a cable into the keyboard of a computer whose screen is off, with a small trophy sitting on the same desk.

A blind replay of one recorded trajectory per task matches or beats the frontier model it was copied from on deterministic benchmarks like OSWorld.

In plain words

  • A tiny replay program matched or beat a leading artificial intelligence system on fixed computer-use tests.
  • The program records one successful series of actions for each task, then repeats those actions without checking the screen.
  • Because the tests do not change, repeating memorized actions can score as well as solving each task anew.
  • The researchers propose varied, checked test setups and measurements that show how uncertain each score is.
  • Better tests would help researchers distinguish adaptable computer control from programs that memorize fixed steps.

Appeared in

Subscribe

Get the brief in your inbox

Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.

  • Weekdays at 8:45am IST, one lead story and 6 to 9 items.
  • Sundays, an argued synthesis rather than a recap.
  • One click to leave, and quiet days say so in the subject line.
How often

Weekdays 8:45am IST + Sundays. Unsubscribe in one click.

You're asking for The Agentic Brief by email at the cadence you picked. You can unsubscribe in one click from any issue, and your address is never sold or shared.