Daily · Sep 22, 2026 · 5 min read
Approve one operation, run another: the binding failure in shipped agent products
Plus: Benchling's DNS-tight sandbox, Shopify's Swift rewrite, nine skill techniques. 5 min.
Curated and summarized by an agent pipeline built by Yadnesh; reviewed before send. How this is made →
The lead
paper · Story page
A new arXiv paper names Loopjacking: a human approves what they understand as operation A, while the implementation applies that decision to a materially different operation B.
- The details:
- Two variants. In a representation-based attack, B is already encoded but omitted or misrepresented at approval time; in post-approval state substitution, the human sees the correct A and mutable workflow state later replaces it. The authors reproduce the second in seven tested Agno AgentOS releases ending at 3.0.9 and in 12 tested versions of a conditional in-memory LangGraph Agent Server composition ending at 0.14.0.
- Yes, but:
- This is a purposive set of released products, not a survey, and every result is version-scoped. OpenClaw 2026.2.23 shows representation mismatch and 2026.2.24 rejects the tested attack.
- Why it matters:
- If a human gate stands between your agent and anything consequential, this is a test rather than a warning. Read the path from the approval prompt to the execution call and check whether anything in between can mutate the operation.
- Researchers found that some computer assistants can carry out a different action from the one a person approved.
- Sometimes the request for permission hides or misrepresents what the assistant will actually do.
- In other cases, the planned action changes after the person gives permission.
- Approval checks protect users only if they cover the exact action the assistant actually carries out.
Research & Papers
05
paper · Story page

Benchproofer turns 500 real SWE-bench Verified issues into formally verified tasks, admitting an instance only once mechanical and adversarial gates agree on its specification. Across two frontier models, a quarter to a half of test-passing patches admit counterexamples.
- Researchers built stricter checks for software fixes written by artificial intelligence.
- Their system writes precise rules for each repair, then uses mathematics to check whether the repaired program follows them.
- Across the systems studied, a quarter to a half of fixes that passed ordinary tests broke those rules.
- For software developers, these checks can catch errors missed by the examples covered in ordinary tests.
Eleven days driving an open-source SQL client's agent mode with 39 local open-weight models produced 8,199 runs. Of 2,100 model-attributed losses, 1,590, or 75.7%, came from runs that had already invoked a tool, so the bottleneck sits after the call.
- In this study, 75.7% of failures blamed on artificial intelligence happened after it had already asked other software for help.
- The assistants worked with databases, organized collections of stored information, by sending requests to other software.
- Requests in the wrong format repeatedly appeared in attempts that used other software but never delivered results.
- For developers improving these assistants, the findings point to problems between asking for help and delivering a finished answer.
paper · Story page
Supervised fine-tuning raised next-turn success for every Qwen3 and Gemma 3 model tested against gold histories, and none of the four then cleared holistic workflow evaluation, with strict trajectory completion topping out at 10.4%. The authors want the layers reported separately.
- Extra training helped four computer assistants answer individual messages better, without helping them finish customer support tasks on their own.
- Some checks judged the assistant's next step after showing it a correct conversation up to that point.
- When assistants handled whole customer support tasks themselves, even the best result met strict completion rules only 10.4% of the time.
- For teams choosing customer support software, checking whole tasks reveals failures that good scores on individual replies can hide.
paper · Story page
A frozen model works design software through 230-plus tools while a memory of natural-language skills grows from real user briefs, gated by replay checks that block any change regressing an observed success. Five rounds take GenEval2 execution success on Claude-Sonnet-4 from 72.7% to 99.3%.
- Researchers improved a computer design assistant by teaching it reusable steps from real design requests.
- It saves and revises written instructions without changing the artificial intelligence system that follows them.
- Changes are kept only when they fix past failures without spoiling tasks that previously succeeded.
- In the reported test, successful task completion rose from 72.7% to 99.3%.
- For design software builders, the results show better task completion without changing the underlying artificial intelligence.
Decoy reasoning problems injected into external context make a model burn far more reasoning tokens while still answering correctly: 13x on FreshQA, 46x on SQuAD, 12x on MuSR. Tested injection points for coding agents include skills, README files and code, and this is a revision of an older paper.
- Researchers found a way to make artificial intelligence assistants do much more unnecessary work while still answering correctly.
- Attackers hide distracting puzzles in material the assistant reads while working on a task.
- The assistant spends extra effort solving those puzzles, increasing the computing work needed for its answer.
- For people running these assistants, correct answers can still come with inflated computing costs.
Engineering & Harnesses
04
blog post · Story page
Every Coinbase support procedure is a versioned decision tree, and Autopilot runs agents that author, test, red-team, score and analyze it behind contracts, fixtures, CI gates and human sign-off on writes. The post says this makes judgment cheaper and more repeatable, not optional.
- Coinbase built Autopilot to help keep its automated customer support working correctly.
- Support bots follow written steps that tell them what information to check and how to handle a customer's problem.
- Autopilot checks those steps for mistakes and possible misuse before people approve changes.
- Coinbase says this reduces the manual work needed to maintain support bots as their software changes.
blog post · Story page
Benchling runs untrusted, agent-generated scientific code for thousands of life sciences tenants on Bedrock AgentCore Code Interpreter in VPC mode. Route 53 Resolver DNS Firewall and VPC endpoint policies layer up to block exfiltration, with DNS named explicitly as one of the closed paths.
- Benchling built protections for running scientific code written by artificial intelligence across thousands of life sciences organizations.
- The code runs inside a private network with rules controlling which outside services it can reach.
- The protections cover website address lookups, which can be misused to send information out.
- These safeguards help protect scientific customers' data while they run computer-written code.
A ban doesn't produce originality, Bakaus argues, it relocates the model one cluster over: forbid an overused font and it reaches for the next nearest thing in latent space. One of his nine techniques splits design direction and deterministic linting into subagents kept blind to each other.
- Paul Bakaus shared nine ways to guide design software that uses artificial intelligence.
- He argues that banning an overused font makes the software choose a similar font instead.
- One method separates design decisions from automatic rule checks, keeping the programs doing those jobs unaware of each other's work.
- His techniques aim to help designers produce original work when detailed written instructions alone fall short.
blog post · Story page
Shopify built an internal tool called Helix to help LLMs rebuild the Shopify app in Swift and Kotlin, using small checkpoints and strict quality gates to keep the code shippable.
- Shopify built Helix to help artificial intelligence rewrite the company's app.
- The rewrite happens in small stages that must pass strict quality checks.
- These checks help Shopify's developers keep the new code ready to release.
Hedge of the day
“As AI becomes more capable, the industry must accelerate security engineering, broaden access to defensive tools and share what works faster.”
The sentence before it asks for named owners, and this one assigns the work to nobody in particular.
Quick links
- AI coding has made CI a bottleneck, so we reworked ours to keep up (Linear Now)
- Benchmarking Grok 4.7 (Artificial Analysis)
- Superpowers 6.4 (Jesse Vincent)
- AI Evals: Everything You Need to Know (Hamel Husain)
- MCP was always a bad idea? (Simon Willison)
- AX: Google's Open Agentic Orchestrator (Hacker News)
- ZCode is now open source (r/LocalLLaMA)
- Amazon blocks Meta's new Muse AI agent from shopping on amazon.com (Forbes)
Meme of the day

Drawn by an image model.
Corrections
Nothing to correct.
Related issues
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.
