Story · Zubin Aysola (Weights & Biases)

How We Built an Agent That Improves Itself (Zubin Aysola (Weights & Biases))

talk · via AI Engineer (talks) · Story page

Weights & Biases turns a production trace into an offline eval task, finds the bug, writes the fix and benchmarks the new agent against the one in production. The research and production agents are kept byte-for-byte identical, and the suite runs 886 tasks including simulated multi-turn users.

In plain words

  • Weights & Biases demonstrated artificial intelligence that helps find and fix problems in its own software.
  • It turns records of real interactions into repeatable tests, including conversations with a computer pretending to be a user.
  • It writes a fix, then compares the revised software with the version people are currently using.
  • The team starts its experiments with an exact copy of the software people use, so tests reflect the real product.
  • These automatically created tests let the team spend more time improving the software instead of writing tests by hand.

Appeared in

Subscribe

Get the brief in your inbox

Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.

  • Weekdays at 8:45am IST, one lead story and 6 to 9 items.
  • Sundays, an argued synthesis rather than a recap.
  • One click to leave, and quiet days say so in the subject line.
How often

Weekdays 8:45am IST + Sundays. Unsubscribe in one click.

You're asking for The Agentic Brief by email at the cadence you picked. You can unsubscribe in one click from any issue, and your address is never sold or shared.