Story · arXiv
Measuring and Exploiting Implicit Trust in LLM Tool-Calling Pipelines (arXiv)
paper · Story page
A framework measures how much a model trusts each of MCP's input channels, then splits prompt injections across two or three of them so that no single channel carries a complete attack.
In plain words
- Researchers tricked some artificial intelligence assistants into giving away private login information by splitting harmful instructions across messages.
- The pieces arrived through descriptions of connected software tools and the messages those tools sent back.
- The pieces seemed harmless alone, but assistants combined them into instructions to send out login details.
- For developers, the tests show that rejecting harmful instructions in one message does not guarantee protection when those instructions are split.
Appeared in
- Split an MCP injection across two channels and resistant models leak at 100%
Sep 18, 2026 · lead story
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.