Story · arXiv
Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection (arXiv)
paper · Story page
A measurement of how much of an injected instruction's authority comes from the reserved chat-template token carrying it, rather than from the text itself.
In plain words
- Researchers found that artificial intelligence systems can be easier to trick depending on how the same malicious instructions are delivered.
- Attackers add fake labels that imitate the labels a system uses to tell who is speaking.
- The same label can arrive as ordinary text or as a built-in signal telling the system how to read the conversation.
- For developers, treating these labels as ordinary text reduced successful attacks across most of the tested kinds of systems.
Appeared in
- A forged chat-template marker loses most of its authority as ordinary subwords
Oct 01, 2026 · lead story
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.