The Agentic BriefNo. 028

Daily · Sep 25, 2026 · 5 min read

Forged control tokens blank an agent's reasoning while the tool call still fires

Plus: a completion gate you can prove, and 8 tools where 83 were. 5 min.

Drawn by an image model.

Curated and summarized by an agent pipeline built by Yadnesh; reviewed before send. How this is made →

The lead

paper · Story page

A new arXiv paper attacks the software around the model. Appending gpt-oss-20b's own channel-control tokens to a user message makes the tokenizer render a reasoning turn as already finished, so the model skips its chain-of-thought and calls the tool.

The details:
Across forty tasks the model already completes, the reasoning channel falls from a mean of 52.5 tokens to zero on every trial while the http.post still fires. Two monitors catch the unsafe request on every plain trial and no forged one, and on overtly malicious requests the attack converts 39.6% of refusals into completed exfiltrations.
Yes, but:
It's one open-weights model under its own published tool sandbox. The paper's claim is that safety is a joint property of the model and the decoding harness, not that other templates parse control tokens this way.
Why it matters:
If your oversight reads the reasoning trace in the tested gpt-oss-20b sandbox, anyone who can put text in the prompt can empty what you're reading while the tool call goes through. Check what your template renderer does with model-reserved tokens.
  • Researchers made an artificial intelligence assistant skip its written reasoning while continuing to take actions.
  • Special formatting in a user's message made the software treat the assistant's reasoning as already complete.
  • With no reasoning to read, two safety checks missed unsafe requests they had previously caught.
  • The findings show developers how data theft can escape safety checks that depend on the assistant's written reasoning.

Research & Papers

05

  • An agent card's name is metadata in A2A, with no collision semantics defined for it. Six of seven pinned open-source integrations still treated it as a local routing identifier, sending a request addressed to a trusted peer to an attacker-controlled one.

    • Researchers found that attackers could impersonate trusted artificial intelligence assistants by copying their names.
    • Some systems used these names to decide where to send requests, even though names were only meant as readable labels.
    • In six tested systems, requests intended for a trusted assistant went to one controlled by an attacker.
    • Users risked sending requests to the wrong recipient, although tests found no direct transfer of the trusted assistant's access details or tools.
  • A bounded loop is a worker, a gate the worker can't write to, and a declared budget. Termination is proved when the repair budget is global rather than per node, and nothing reaches DONE without a gate verdict in an append-only hash-chained ledger.

    • Researchers designed a system to keep artificial intelligence assistants from running or spending without limits.
    • A separate checker the assistant cannot change must approve each step before it counts as finished.
    • A repair limit covering the whole job guarantees that repeated attempts eventually stop.
    • People running these assistants get spending limits that apply even while an attempt is underway.
  • Only 78 of the 125 all-fail tasks in a frozen Terminal-Bench 3 and Frontier-Bench 0.1 record survive as certified-unsolved candidates. The rest involve broken oracles, infrastructure failures, verifier bypasses, or tasks with insufficient evidence of solvability.

    • Researchers found that many failed artificial intelligence tests had problems with the tests themselves.
    • They checked whether the tasks could be solved and whether the systems judging answers worked properly.
    • Only 78 of 125 tasks with no successful solution appeared to be fair tests that the assistants still could not pass.
    • People comparing artificial intelligence systems could mistake a broken test for evidence that an assistant lacks ability.
  • 6,774 merged pull requests from Codex, Copilot, Devin, Cursor and Claude Code, followed into their later fixes against 5,044 human pull requests from the same repositories. The agent ones draw verified follow-up fixes at 1.62 times the odds.

    • A study found that software changes from artificial intelligence assistants received later fixes more often than changes from people.
    • Researchers compared 6,774 changes from assistants with 5,044 from people after those changes were added to the same software projects.
    • Human reviewers and an artificial intelligence system checked which later changes actually repaired that earlier work.
    • For software teams, adding changes from these assistants to a project may still leave repair work to do.
  • Agents propose JSON Patch mutations against shared structured state instead of talking to each other, and a deterministic kernel validates each one before it commits. On 630 matched ALFWorld episodes: 84.6% success against 30.8% for LangGraph and 61.6% for Flock.

    • PatchBoard lets artificial intelligence assistants coordinate their work through changes to a shared record.
    • A planning assistant sets the structure of the shared record and the rules for doing the task.
    • A separate program checks every proposed change against those rules and each assistant's permissions before saving it.
    • In the reported tests, assistants using PatchBoard completed 84.6% of tasks successfully, more than either comparison system.

Engineering & Harnesses

03

  • Backslash took the 8,000 most-starred MCP servers on GitHub and found at least one risk finding in 29% of them, with confirmed unintended remote code execution in six.

    • Backslash found security risks in 29% of 8,000 popular programs that connect artificial intelligence assistants to outside tools.
    • In six cases, someone could make the computer running the program carry out unwanted instructions from elsewhere.
    • The report offers security advice for developers deciding how to connect assistants to other software.
  • Zenity Labs reached Agentforce account data through a public Web-to-Lead form, with no login and no victim click, and got it out over a DNS query. It reports bypassing the guardrails and the Trusted URLs redaction layer; Salesforce has remediated.

    • Zenity Labs found a way to steal Salesforce account information without logging in or requiring the account owner to click anything.
    • The attack put malicious instructions in a public form for potential customers that Salesforce's Agentforce assistant could read.
    • The attack sent stolen information through a request normally used to find a website's address, bypassing security filters.
    • Salesforce fixed the flaw to protect customers' account information from this attack.
  • Notte's MCP server was generated from its OpenAPI schema, one tool per route, and had reached 83. Eight hand-written tools built around the browser loop replaced it on September 4, cutting per-turn tool definitions from about 38,000 tokens to about 4,500.

    • Notte cut the number of tools offered to automated browser assistants from 83 to eight.
    • Each new tool covers a complete user task, instead of giving the assistant a separate control for every small software operation.
    • The full collection remains available separately for developers who need it.
    • Developers using the shorter list give their assistants much less instruction text to read each time they respond.

Hedge of the day

“No unsupported value detected by recorded controls remained in the engraved matrix.”

TerriScan: An Incident-Evaluated, Doctrine-Governed Multi-Agent LLM System for Recalculable Urban Indicator Production in the Global South (Semantic Scholar)

The values its controls never detected sit outside the claim.


Meme of the day

The brick front of a building at pavement level. A heavy steel door carries a plate reading AUTHORISED ACCESS ONLY, with a camera above it. Beside it, a comment box on a post under a HOW ARE WE DOING sign has a tube running from it into an overfull sack of envelopes on the pavement.
The door passed its annual penetration test. The box is not in scope. More in the hall of fame

Drawn by an image model.

Corrections

Nothing to correct.

Related issues

Subscribe

Get the brief in your inbox

Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.

  • Weekdays at 8:45am IST, one lead story and 6 to 9 items.
  • Sundays, an argued synthesis rather than a recap.
  • One click to leave, and quiet days say so in the subject line.
How often

Weekdays 8:45am IST + Sundays. Unsubscribe in one click.

You're asking for The Agentic Brief by email at the cadence you picked. You can unsubscribe in one click from any issue, and your address is never sold or shared.