Daily · Oct 05, 2026 · 5 min read
VS Code rebuilt its release process around agents and now ships weekly
Plus: an agent that writes its own tools, and a support bot at 80%. 5 min.
Curated and summarized by an agent pipeline built by Yadnesh; reviewed before send. How this is made →
The lead
talk · Story page
VS Code now ships weekly after ten years of monthly releases. Harald Kirschner's account of how is almost entirely about the system around the agents.
- The details:
- The rebuild came in three stages: agent-ready codebases and skills plus TypeScript Go's 10x faster builds; mandatory AI code review, AI issue triage and error stacks that become auto-fix pull requests; then VSC-Bench evals and smaller squads. Agent-written code that gets committed rose from 55% to 86%.
- Yes, but:
- These are Microsoft's numbers about Microsoft's product, from a talk, and code survival counts what got committed, not what held up. Kirschner also names what broke: CI that compounds when 20 agents hit it at once.
- Why it matters:
- If you own the build system, you pay for this first. Kirschner describes slow CI compounding when 20 agents hit it at once, alongside TypeScript Go's 10x faster builds and mandatory AI code review.
- Microsoft now updates its app for writing computer programs every week instead of every month.
- Artificial intelligence (AI) can operate the app to check and correct the computer instructions it writes.
- The team requires AI to check proposed changes to the app.
- These changes helped a small team deliver frequent updates to more than 50 million users.
Research & Papers
03
paper · Story page
An evidence-informed review of the MCP security surface: trust boundaries, tool poisoning, privilege propagation, credential use, cross-tool exfiltration and supply-chain compromise. It leans on empirical results from 2025 and 2026.
- Researchers reviewed how connecting artificial intelligence (AI) to outside tools can create security risks.
- These connections let AI use information and take actions in other services.
- The review examines how misleading tool descriptions can affect the AI's decisions.
- Developers using these connections face risks of stolen information or actions taken without proper permission.
paper · Story page
The authors split a differential-testing pipeline into Agent Skills for program generation and feedback-guided mutation, then compare four configurations. Feedback-guided mutation gave the largest gain at 12 percentage points; adding more procedural knowledge cost 3.
- Researchers split artificial intelligence (AI) software testing into separate instructions for creating tests and improving them.
- The tests look for programs giving different answers to the same calculation.
- Changing tests based on earlier results raised the rate of finding conflicting answers by 12 percentage points.
- For software testers, the study suggests that more detailed instructions do not always help uncover problems.
paper · Story page
Data-reflector confines the LLM to reading intent, picking tools and parsing filters, and routes every number through fixed functions running locally. Six participants ran it on six clinical and nonclinical datasets, scored partly on whether a second operator could reproduce the run.
- Researchers built Data-reflector, a tool for exploring data through questions written in everyday language.
- Artificial intelligence (AI) interprets each question to choose the right calculation tools.
- Preset computer instructions carry out every calculation.
- A trial with 6 participants checked accuracy and whether another person could repeat the work.
- The tool is intended to make data analysis more accessible to researchers who cannot write computer programs.
Engineering & Harnesses
05
Gemini's Interactions API moves conversation state server-side behind interaction IDs, so you stop hand-managing thought signatures, and adds strongly typed multimodal outputs. Managed Agents run the Antigravity harness in a persistent remote sandbox with loadable sources and a credential-injecting proxy.
- Google demonstrated new tools for building artificial intelligence (AI) helpers that can carry out tasks.
- The service stores conversation history so developers do not have to keep track of it themselves.
- Each helper gets a separate workspace on Google's computers that stays available between tasks.
- Developers have less conversation history and computer setup to manage when building helpers that work over several steps.
Two Temporal primitives, a wait condition and a signal, let an agent pause for a human for minutes or weeks without blocking. Warrick kills the worker mid-approval in a Google ADK and LangGraph demo, then brings it back with no lost work.
- Temporal showed how artificial intelligence (AI) can wait for a person's approval without losing its work.
- The system remembers where a task paused and continues when the person's reply arrives.
- In the demonstration, the program running the task stopped and restarted without losing completed work.
- People can take minutes or weeks to respond without forcing the task to start over.
blog post · Story page
Willison wants pay-by-usage services to ship hard caps by default: after $X a month, cut the thing off and return errors. A warning email at midnight leaves the rogue service billing you while you sleep.
- Willison wants services that charge for each use to stop when a customer reaches their spending limit.
- Programs created with artificial intelligence can generate charges by using paid services.
- Warning emails leave the service running, so bills can keep growing before anyone notices.
- Automatic shutoffs would protect customers from bills above their chosen limits.
talk · Story page
AssemblyAI swapped a bot that resolved 10% of conversations for Joey, which the team reports resolves 80% end to end for about $700 a month. It's the Claude Agent SDK over docs checked out as local Markdown, plus embeddings and one large CLAUDE.md.
- AssemblyAI replaced its old automated customer support helper with a new one called Joey.
- Joey searches saved help documents for information relevant to customers' questions.
- The team demonstrated Joey listening and speaking during a live conversation.
- AssemblyAI says Joey handles 80% of customer conversations without human help for about $700 a month.
Subramani's Strands agent starts with zero tools and writes the ones it needs at runtime, off a system prompt plus three primitives: editor, shell and load_tool. She also covers the sandboxing, constrained permissions and observability you'd want first.
- Sandhya Subramani demonstrated an artificial intelligence assistant that writes extra computer programs when a task requires them.
- It creates a program and starts using it immediately, without restarting.
- She described limiting what it can access and tracking its actions before trusting it with real work.
- Developers could let assistants tackle new tasks without having to build every tool ahead of time.
Hedge of the day
“An exploratory evaluation with agent-native execution reveals scalability challenges for high-iteration Skill workflows.”
The paper counts its improvements in percentage points and its problems in adjectives.
Quick links
- From manual negotiation to automated scheduling: How AI21 manages its GPU fleet with Kueue (AI21 Labs Blog)
- Agents don't need memory, they need documentation (liao.gg)
- AI Coding Agents Are Breaking Big Codebases (Dan Adler (Sourcegraph))
- Strands Decider: Why Not an Encoder? (Marc Brooker)
- The ultimate guide to multi-harness RL (r/LocalLLaMA)
- How we cut time to first byte by 94% for long prompts (LiteLLM (BerriAI))
- I built a code knowledge graph tool that's actually MIT licensed (fully local, no cloud) (r/LocalLLaMA)
- AI Gateway Guardrails in 2026: How to Secure LiteLLM, Kong, TrueFoundry and Agent Router (and 12 Guardrail Providers Compared) (Pillar Security Blog)
- Unitree just dropped UnifoLM-WLA-1.0: a single 6B model that does 64 whole-body + tabletop tasks on a real humanoid (r/LocalLLaMA)
- The Rise of Overfit Inference Engines (r/LocalLLaMA)
Meme of the day

Drawn by an image model.
Corrections
Nothing to correct.
Related issues
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.