Daily · Aug 10, 2026 · 5 min read

An independent harness matches DeepSeek's 82.7%, and agents outrun their verifiers

Plus: 4 talks from AI Engineer, a GitHub retirement, and sycophancy for smart people. 5 min.

Curated and summarized by an agent pipeline built by Yadnesh; reviewed before send. How this is made →

The lead

Reddit thread · Story page

What happened:
An independent evaluator reproduced DeepSeek's reported 82.7% score on Terminal-Bench 2.1 for V4 Flash 0731, using the public, downloadable Ante harness rather than DeepSeek's own unreleased "minimal mode" setup.
The details:
The run logged 368 successful trials out of 445 across 89 tasks, five trials per task, at max reasoning effort through OpenRouter. The full Harbor job is public, down to the pinned configuration and per-trial rewards, exceptions, durations, and token usage.
Yes, but:
The replicator is Ante's own author, it's a single run with a ±1.79 standard error, and the author notes the model seems sensitive to its harness. The number that got matched came from a harness DeepSeek hasn't released yet.
Why it matters:
If you pick models off Terminal-Bench scores, this one now comes with a pinned config you can re-run instead of a vendor claim you can't. The harness-sensitivity note cuts the other way: the score may travel with the harness, so your own agent stack may not see 82.7%.

Research & Papers

Engineering & Harnesses

  • talk · Story page

    Chatterjee argues AI-written code leaves verification debt that human review alone can't reliably contain: a Wharton study he cites suggests reviewers followed AI advice nearly 80% of the time even when it was instructed to lie confidently. His fix is zero-trust, multilayer verification that checks generated code by methods independent of the model that wrote it.

  • talk · Story page

    Singh's meeting bot sat in a Google Meet for four hours, heard a passerby wish coding agents had acceptance criteria, opened the ticket itself, and prototyped the change. The narrower claim underneath: customer calls and team meetings now yield prototyped ideas and usually a few shippable pull requests, with one agent session reachable from Slack, desktop, and GitHub at once.

  • talk · Story page

    Resolve AI's background agents notice a release tag dropped in Slack, read what changed, and write a monitoring plan for that release alone, deciding on their own when to check back. The target is everything that routes around CI/CD: feature flags and infrastructure changes that ship with no monitoring at all.

  • talk · Story page

    Dailey names the gap between agent-amplified output and actual impact "velocity sickness": unread content, unmergeable pull requests, mornings spent declaring agent bankruptcy. His fix is keeping critical decisions with humans in a durable document layer, separate from the ephemeral chats where implementation happens.

Product & Releases

  • blog post · Story page

    GitHub retired GitHub Models, the playground and unified model API that let Actions workflows run prompts on the GitHub API key already in the environment. Willison migrated his repository-summary workflow to OpenAI and bets the shutdown fits a pattern: coding-agent usage made free or subsidized tokens prohibitively expensive to offer.

Community

  • blog post · Story page

    Goedecke questions whether frontier models have gotten less sycophantic or merely better at flattering smart, neurotic information workers: the effective move is disagreeing with you without making you feel stupid. Worth holding in mind anywhere you let a model grade your work.

From X

Corrections

Nothing to correct.

Subscribe

Get the brief in your inbox

Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.

  • Weekdays at 8:45am IST, one lead story and 6 to 9 items.
  • Sundays, an argued synthesis rather than a recap.
  • One click to leave, and quiet days say so in the subject line.
How often

Weekdays 8:45am IST + Sundays. Unsubscribe in one click.

You're asking for The Agentic Brief by email at the cadence you picked. You can unsubscribe in one click from any issue, and your address is never sold or shared.