Daily · Sep 21, 2026 · 5 min read
Giving an agent the decoder's confidence never beat a plain deterministic gate
Plus: three agent architectures priced against the same task, and nine quick links. 5 min.
Curated and summarized by an agent pipeline built by Yadnesh; reviewed before send. How this is made →
The lead
paper · Story page
Researchers tested whether telling an agent how confident an upstream decoder was makes its actions safer, using 1,065 brain-computer-interface episodes from 47 people with ALS and five language models.
- The details:
- At matched coverage, no confidence-prompted arm pushed unfaithful execution below what a deterministic gate achieved, and two of the five models did significantly worse. A follow-up gave ten models the same command vocabulary as the deterministic resolver, and no direct agent arm beat the resolver's risk-coverage frontier.
- Yes, but:
- The apparent safety gains, up to 22 percentage points, came from acting less often, and sometimes the abstention was an invalid tool call rather than a refusal. A model that fails to emit a valid action looks careful in the numbers.
- Why it matters:
- In this brain-computer-interface study, decoder confidence in the prompt did not reduce unfaithful execution below a deterministic gate at matched coverage. The hybrid architecture kept admission authority outside the model and let it propose corrections, which extended coverage in five of ten models without observed unfaithful executions.
- Researchers tested safer ways for artificial intelligence (AI) to act on commands from brain signals.
- Telling AI how certain each command was did not reduce mistakes compared with fixed approval rules when both acted equally often.
- A follow-up let AI suggest corrections, while a separate check decided which commands could be carried out.
- For users of brain-controlled computers, this combination handled more commands in five of ten tested systems, without observed mistakes in carrying them out.
Research & Papers
03
paper · Story page

MCP task states are coarse and self-reported, so a worker thread that deadlocks keeps showing up as working. ReliHarness probes the thread itself with a heartbeat timeout and an operating-system liveness check, reporting 100% detection coverage and zero false positives on the crash and deadlock scenarios tested.
- Researchers built ReliHarness to detect when software used by artificial intelligence (AI) stops working.
- A task can still appear to be running after the software carrying it out has frozen or stopped.
- ReliHarness checks for regular signals from that software and asks the computer whether it is still running.
- It detected every tested failure without wrongly flagging working software.
- For people overseeing long tasks, these checks can reveal failures that ordinary progress reports miss.
paper · Story page
One calibration workflow organised three ways: instructions consulted on demand, tool interfaces supplied in advance, one specialised agent per stage. All three produced models of equivalent accuracy, so the choice is cost and endurance, and the cheapest option got unreliable once the task ran long.
- Researchers compared ways for artificial intelligence (AI) to adjust computer simulations of how water drains through cities.
- The AI followed instructions as needed, used prepared tools, or split the work among systems handling different stages.
- All approaches produced equally accurate simulations, but the cheapest became unreliable when tasks grew longer.
- For people planning city drainage, choosing an approach changes the cost and reliability of the work.
paper · Story page
Twelve GPT-4o-mini sellers and twelve buyers over twenty rounds with hidden quality. A twenty-replicate batch put the forum-associated difference in unseeded sellers making at least one false quality claim at 0.850 under private history and 0.483 under public history, counting observable overstatement only.
- Researchers linked seller-only chats to more exaggerated product claims in an artificial intelligence (AI) shopping simulation.
- The increase was smaller when everyone could see the market's past activity.
- The findings only cover markets including sellers given special experimental treatment, with some results changing when product descriptions were reworded.
- The study counted exaggerated claims without assuming the AI intended to deceive.
- For people designing automated markets, the results suggest that what sellers can see and discuss may affect product claims.
Engineering & Harnesses
04
talk · Story page
Ask a harness for 200 queries a second and it may deliver 38, then print the results as though it ran 200. The global interpreter lock caps a single-process client, one shared 20% throughput win ran at temperature zero, and Inference Perf logs scheduled against actual send times.
- A presentation explains why speed tests can give misleading results about artificial intelligence (AI) services.
- Software sending test questions can fall behind, making a working service appear slow.
- The presenters' testing tool records when questions should be sent and when they actually leave.
- People running AI services can use those records to check whether apparent delays come from the testing software.
Imprint's year of agent adoption reads as workflow changes, not model changes: about ten local workspaces, each with an independent checkout of every repository, so agents can generate pull requests across frontend, backend, infrastructure and data. Then the company left Jira for Linear over task visibility and permissions.
- Imprint changed how its staff organize software work as they began using artificial intelligence (AI) more widely.
- Separate computer workspaces each contain all its software projects, so AI can prepare connected changes across those projects.
- The company moved everyone's tasks to Linear to make work easier to see and access rules less complicated.
- For Imprint's staff, these changes address the difficulty of coordinating AI work across software projects.
Black-box probing suggests Instinct keeps memory as git-tracked Markdown files found with grep rather than vectors, so read it as a reconstruction and not confirmed internals. The post rebuilds the same design on Supermemory in about 60 lines.
- A post suggests Instinct remembers information by saving it in text files.
- The reconstructed design keeps a record of changes to those files.
- It finds saved details by looking for matching words in the files.
- Developers can follow the post to build a similar way for software to remember information.
One user reports running Qwen 3.8 27B for about 21 days on one RTX 3090 to build a CUDA inference engine, on roughly 12 human messages, getting working kernels that never beat llama.cpp. Compaction ate about 83 hours, and the same card had to host the agents and run the engine under test.
- One user reports letting an artificial intelligence system write software for about 21 days.
- Its task was to build software that could run artificial intelligence more efficiently on the same computer.
- Written rules told it when to ask the user for help.
- The system had to stop running so its new software could be tested on the same hardware.
- For this user, days of largely independent work produced working software that was no faster than an existing alternative.
Product & Releases
01
Von is an open-source 395M System One model pitched as a drop-in replacement for TypeSafe's JEV. Its author says it runs on CPU in 1 to 2 GB, answers in 25 to 300 ms, and beats JEV in all benchmarks.
- Von's creator released an artificial intelligence system with code that others can inspect and change.
- The creator says it can run entirely on the main chip that carries out a computer's general tasks.
- The creator reports better results than TypeSafe's existing system in every comparison test they ran.
- Developers using TypeSafe's system could switch to Von without changing how their software uses it, according to its creator.
Quick links
- Can Jev Be a Better Agent Evaluator? (LangChain)
- llm-keys-ui 0.1 (Simon Willison)
- Quoting voxium (Simon Willison)
- What's New in Inference Engineering (Philip Kiely (Baseten))
- Weight Folding, CUDA Streams, and the Bug That Made My Model Speak Backwards (Filip Makraduli)
- Brownfield Agentic Engineering (Addy Osmani)
- FunPilot: Runtime Performance Diagnosis and Remediation for Serverless Applications with LLMs (Semantic Scholar)
- FinDS-Agent: A Cloud-Edge Collaborative Data Science Agent for Financial Analytics (Semantic Scholar)
- System One models like Jev can train their own replacements (Sean Goedecke)
Meme of the day

Drawn by an image model.
Corrections
Nothing to correct.
Related issues
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.
