Story · arXiv

An Empirical Cost Attribution of Context-Compression Gateways in Multi-Turn Coding Agents (arXiv)

paper · Story page

A two-axis chart over a row of tick marks: one straight diagonal line rising steadily from the origin, and a second line starting flat and curving upward to overtake it near the right edge.

A team instrumented Paritok, a production compression gateway, between coding agents and frontier models, then split the token bill three ways. Only tool-schema filtering, which strips a fixed block of roughly 21K to 57K tokens on a typical turn, comes out reproducibly positive; content compression returns about 2% of the cache-priced prefix per turn.

In plain words

  • Researchers compared ways to reduce the running costs of artificial intelligence assistants that write code.
  • One approach removes some descriptions telling the assistant how to use its available software tools.
  • Another shortens the file contents shown to the assistant, with savings growing as that text is reused in later exchanges.
  • For people paying for coding assistants, removing tool descriptions was the only approach that consistently saved money in these tests.

Appeared in

Subscribe

Get the brief in your inbox

Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.

  • Weekdays at 8:45am IST, one lead story and 6 to 9 items.
  • Sundays, an argued synthesis rather than a recap.
  • One click to leave, and quiet days say so in the subject line.
How often

Weekdays 8:45am IST + Sundays. Unsubscribe in one click.

You're asking for The Agentic Brief by email at the cadence you picked. You can unsubscribe in one click from any issue, and your address is never sold or shared.