Topic · 21 stories
Infra & sandboxes
Stories
Homa: The End of TCP for AI Clusters (John Ousterhout (Stanford), AI Engineer)
talk · Story page
Why small coordination messages queue behind big transfers
Qwen3.8 Flash Next now at 1.2k t/s prefill on Strix Halo (r/LocalLLaMA)
Reddit thread · Story page
One author's llama.cpp optimizations for Qwen3.8 Flash Next on Strix Halo
Moli Builds a Rust Browser That Uses 10x Less Memory Than Chrome (AlphaSignal)
newsletter item · Story page
Rust headless browser that skips rendering, per AlphaSignal
RASER: Resilient Agent Scheduling and Execution Runtime for HPC Clusters (Semantic Scholar)
paper · Story page
RASER runs agent workflows on Slurm with work stealing through shared-filesystem queues, application-level checkpointing paired with Slurm requeue, and Apptainer isolation without image changes. It reports nearly 39% lower makespan than static partitioning at near-full CPU utilization, and it tests recovery after simulated preemption.
vLLM x AgentX: Optimizing for Real-World Agentic Serving (vLLM Blog)
vLLM walks through KV cache management, parallelism, scheduling and prefill/decode disaggregation for agent traffic, and reports up to 130K tokens per GPU-second on SemiAnalysis AgentX plus a 14.6x to 106x serving-cost advantage over Opus 5. Those are the project's own numbers.
Optimizing delta weight syncs for managed rollouts (Baseten)
Baseten syncs RL weight deltas across clusters in under 40 seconds
Latency-Aware Orchestration for Multi-Agent LLM Workflows on Heterogeneous GPUs (Semantic Scholar)
paper · Story page
Prediction-guided scheduling for multi-agent workflows on mixed GPU pools
Same Request, Different Answer: Quantization Amplifies Cache-Induced Divergence in LLM Serving (arXiv)
paper · Story page
An eighty-episode agentic tool-use workload measures how prefix caching changes trajectories under fixed seeds, across two engines and four weight formats.
Cursor Cloud Agents can now run in Vercel Sandbox (Vercel)
blog post · Story page
Cursor keeps the agent harness and inference loop, and Vercel Sandbox supplies an isolated Firecracker microVM per agent request, with Vercel Functions and Workflow as the control plane that claims queued requests, provisions workers, monitors sessions, and cleans up. You get a scale-to-zero worker pool, durable retries, and short-lived user-scoped credentials, and Self-Hosted Machines requires a Cursor Enterprise plan.
Introducing Sandbox Computer Use for Mastra Agents (Mastra)
Mastra agents can now drive a sandboxed Linux desktop through E2B Desktop or Daytona, with 11 computer tools for browsing, forms, downloads, screenshots, and terminal programs. A SandboxComputer interface exposes screen size, cursor position, and a stream URL so you can watch the agent work.
Set per-user budgets on AI Gateway (Vercel)
blog post · Story page
per-user spending limits on Vercel AI Gateway
The Half Life of Agent Infrastructure (Ben Kus (Box))
talk · Story page
Agent infrastructure's half life, measured in months
vLLM v0.28.0 (Hacker News)
HN thread · 106 points on HN · Story page
The vLLM project ships version 0.28.0
FinOps for AI Agents: Who Spent All the Tokens? (Tisha Chawla & Susheem Koul (Microsoft))
talk · Story page
Microsoft's agent cost control plane, steering runs instead of killing them
I hosted Kimi K3 (2.8T parameters) using 8 B300s. 92 tok/s, $190 per million tokens (r/LocalLLaMA)
Reddit thread · Story page
Real serving numbers for a 2.8T-parameter model on eight B300s
From LLM Inference to Agentic Workloads: Characterization and Implications for Serving Systems (arXiv)
paper · Story page
AgentSysBench: where latency and memory land across ten agentic apps
Docker Sandboxes – Disposable, isolated sandboxes for AI agents (Hacker News)
HN thread · 375 points on HN · Story page
Disposable, isolated sandboxes for agents; 375 points on Hacker News.
Smart Routing in Unity AI Gateway: Match frontier quality with 30%+ lower cost per task (Databricks)
blog post · Story page
Databricks says smart routing in Unity AI Gateway spreads coding tasks across a pool of models and harnesses, matching frontier quality at more than 30% lower cost per task.
Monitor on-premises and multi-cloud AI agents with AgentCore Observability (AWS ML Blog)
blog post · Story page
AgentCore observability for agents running outside AWS
Databricks Acquires Electric to Give Every AI Agent its Own Postgres (AlphaSignal)
newsletter item · Story page
Databricks is acquiring ElectricSQL to embed a full Postgres inside every AI-agent sandbox, syncing local state back to its Lakebase platform.
Always-on agents run production without the on-call tax (Justin Smith (Resolve AI))
talk · Story page
Resolve AI's background agents notice a release tag dropped in Slack, read what changed, and write a monitoring plan for that release alone, deciding on their own when to check back. The target is everything that routes around CI/CD: feature flags and infrastructure changes that ship with no monitoring at all.
Issues that covered it
- Daily
Split an MCP injection across two channels and resistant models leak at 100%
Plus: reward hacking at 57.2% of rollouts, and Exa searches the web by date. 5 min.
Sep 18, 2026 · 5 min

- Daily
An agent deleted an AML control, and benchmark scaffolds do the model's work
Plus: OpenAI's 10,000-agent proof, a 297-iteration model loop, MCP merges Skills. 6 min.
Sep 14, 2026 · 6 min

- Daily
Per-phase model routing cuts agent cost, RASER on Slurm, Amp steers mid-run
Plus: 2 papers, a context-mode split, and OpenAI's Navier-Stokes claim. 5 min.
Sep 09, 2026 · 5 min

- Daily
Prefix caching changes agent runs, and cross-family reviewers beat self-review
Plus: a harness that moves solve rates 4x, and 883 commits to nowhere. 6 min.
Sep 08, 2026 · 6 min

- Daily
GPT-6 Astra's score hinges on its harness, and agents rot geometrically with each step
Plus: Cursor agents in Vercel microVMs, Composio's six missing primitives, ACLE-MCP. 5 min.
Sep 04, 2026 · 5 min

- Daily
Google's Mantis bug-fixing harness, and privilege escalation in 12 agent harnesses
Plus: Cline's 11M-user rollout, ContextPipe, and Claude driving your desktop. 5 min.
Sep 03, 2026 · 5 min

- Daily
Realistic prompts drop coding-agent scores, and tool filtering beats prompt rules
Plus: 7 papers, 2 releases, one empty-handed mugger. 5 min.
Sep 01, 2026 · 5 min

- Daily
Maersk's 100,000 corrections, preference-trap evals, and Tencent's Hy4 Preview
Plus: guardrails at Navan, a one-way fiber, and the benchmark card that names its scaffold. 5 min.
Aug 31, 2026 · 5 min

- Daily
Poisoned agent memory beats screening, and prompt rules keep failing as boundaries
Plus: bash-sed-cat at AIDAChip, a breach in Thailand, 8 quick links. 5 min.
Aug 25, 2026 · 5 min

- Daily
Aborted agent branches live on in the KV cache, plus StateM's harness scaling
Plus: coherence debt, the recall trap, Cline's open-weight evals. 5 min.
Aug 19, 2026 · 5 min

Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.









