Story · arXiv
An Empirical Cost Attribution of Context-Compression Gateways in Multi-Turn Coding Agents (arXiv)
paper · Story page

A team instrumented Paritok, a production compression gateway, between coding agents and frontier models, then split the token bill three ways. Only tool-schema filtering, which strips a fixed block of roughly 21K to 57K tokens on a typical turn, comes out reproducibly positive; content compression returns about 2% of the cache-priced prefix per turn.
In plain words
- Researchers compared ways to reduce the running costs of artificial intelligence assistants that write code.
- One approach removes some descriptions telling the assistant how to use its available software tools.
- Another shortens the file contents shown to the assistant, with savings growing as that text is reused in later exchanges.
- For people paying for coding assistants, removing tool descriptions was the only approach that consistently saved money in these tests.
Appeared in
- Dormant prompt injections land on nine production agents where direct orders fail
Sep 23, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.
