The Agentic BriefNo. 027

Daily · Sep 24, 2026 · 5 min read

Two MemOS packages shipped credential stealers into the agent memory layer

Plus: an agent that spent eight days rewriting itself, and an MCP hijack that rides tool metadata. 5 min.

Drawn by an image model.

Curated and summarized by an agent pipeline built by Yadnesh; reviewed before send. How this is made →

The lead

blog post · Story page

Socket reports that two packages from MemTensor's MemOS, an open-source memory framework for LLMs and AI agents, were compromised: the npm plugin @memtensor/memos-cloud-openclaw-plugin and the PyPI package MemoryOS.

The details:
Both drop cross-platform Go binaries that exfiltrate developer secrets. That's compiled code, not an obfuscated postinstall snippet, and it runs the same on a laptop or in CI.
Yes, but:
Pinning back to a clean version isn't enough. If those binaries ran, treat reachable credentials as compromised and rotate them.
Why it matters:
MemOS is the sort of dependency you add once and stop thinking about. Anyone who loaded an affected version should rotate reachable credentials as well as pinning to a clean release.
  • Two tools that help artificial intelligence remember information were tampered with to steal developers' secrets.
  • The affected software, MemoryOS and @memtensor/memos-cloud-openclaw-plugin, installs extra programs that send developers' private information elsewhere.
  • Developers using either tool risk exposing secrets across different types of computers.

Research & Papers

04

  • Runtime policies are natural-language instructions and action denials the harness applies at the states that preceded observed failures, leaving weights and the user prompt untouched. Across the 87-task Terminal-Bench 2.1 suite, repeated success for GPT-5.6 Sol moves from 64.4% to 73.6%.

    • Researchers made an artificial intelligence assistant more consistent at finishing tasks by giving it rules based on earlier mistakes.
    • When the assistant reaches a situation linked to past failures, the software gives extra instructions or blocks particular actions.
    • The approach needs no retraining of the assistant or changes to the user's request.
    • In tests, one assistant completed both tries on 73.6% of tasks, up from 64.4%, suggesting more dependable help for users.
  • AIDE^2 proposes changes to its own research-agent code, benchmarks the modified versions of itself on AI R&D tasks, and keeps whatever wins on hidden evaluations. An autonomous eight-day run found seven successive improvements, and the gains carried to four held-out benchmarks.

    • Researchers built an artificial intelligence research assistant that improves itself by rewriting its own software.
    • It tries changes on research tasks and keeps the versions that perform best in hidden tests.
    • Each improved version becomes the starting point for the next round of changes.
    • The improvements carried over to four separate test sets, suggesting the assistant could help researchers with tasks beyond those used during development.
  • A2M optimizes MCP tool metadata so the agent reaches for the attacker's server, then uses execution traces to refine that tool's returns. On LiveMCPBench against GLM-4.6: 93.6% malicious invocation and 74.4% mean attack success, both lower when transferred to other models.

    • Researchers demonstrated a way to trick artificial intelligence assistants through tools offered by outside providers.
    • The attacker rewrites a tool's description to make the assistant more likely to choose it.
    • The attacker then studies records of the assistant's actions to refine harmful instructions sent back by the tool.
    • Users could have information stolen, though the attacks worked less well when tried on other assistants without further adjustments.
  • Two agents each write a patch that passes alone, then the pair breaks when combined. Interference hit 97% of runs on constructed tasks using 12 real Django helpers, and one of 834 runs on mined pull-request pairs; the authors say the constructed rate estimates nothing about practice.

    • Researchers found that changes written by separate artificial intelligence coding assistants can work individually but break when put together.
    • One assistant can change a rule that the other assistant's work still depends on.
    • These conflicts were common in specially constructed tests but appeared only once in 834 runs using previously accepted changes.
    • Teams using coding assistants still need to check combined work, though these artificial tests do not reveal how common conflicts are.

Engineering & Harnesses

03

  • Cursor publishes where its production agent inference spend lands, then describes changes to request assembly, context reuse, and how work gets divided across agents, for a 7% cut in users' token costs without reducing agent quality. Tool outputs and tool-call arguments carry most of the bill.

    • Cursor says it made its artificial intelligence coding assistant cheaper to run without reducing the quality of its work.
    • The software changed how it reuses information the assistant has already received.
    • It also changed when work is divided among several assistants.
    • Users pay 7% less for the text the assistant processes, according to the company.
  • Pass@1 is flat across three generations on one expert-created terminal-bench style task set: 61.5% for Fable 5.1, 60.7% for Opus 5, and 60.7% for Opus 5.5. Pass@5 separates them, and Opus 5.5's 76.7% sits below Opus 5's 79.3%.

    • Snorkel found that newer coding assistants did not consistently do better on its programming tests.
    • The tests checked success on the first attempt and whether allowing five attempts helped.
    • Some Fable 5.1 failures came from stopping early or being unable to recover after errors.
    • People choosing coding help can use these failure details to compare assistants beyond their overall success rates.
  • Anthropic says Claude autonomously found an enzyme system built on a known reverse transcriptase from a jumbo phage, paired with an array of non-coding DNA repeats and an accessory protein. Nobody knows yet what the system does.

    • Anthropic says its artificial intelligence program Claude independently found a previously unrecognized arrangement of biological parts.
    • The system combines a protein that copies genetic material with repeated stretches of genetic material and another protein.
    • The copying protein was already known, but Claude appears to have first noticed the surrounding arrangement.
    • Researchers now have a new system to investigate, though nobody yet knows what it does.

Product & Releases

01

  • The GitHub Copilot app adds configurable local sandboxing that limits which files, network resources, and credentials a command can reach on your machine.

    • GitHub's Copilot app now lets users limit its access when it carries out instructions on their computer.
    • Users can restrict access to files, network connections, and details used to sign into accounts.
    • These limits can reduce the damage if the app carries out an instruction the user did not intend.

Community

01

  • One person who sells automation to small companies says the supplier-order agent they shipped has looked flawless for four months because it gets stuck about twice a week and they fix it first. How much of your uptime is the agent, and how much is you?

    • A seller says a customer's program for ordering supplies gets stuck about twice a week.
    • They fix it before the customer notices, making four months of operation look trouble-free.
    • The seller says fixing the program still takes less work than handling those orders by hand.

Meme of the day

A corridor wall. A clean workplace safety board reads DAYS SINCE LAST INCIDENT above three number cards showing 121. Directly below the board is a square metal access panel with a screw at each corner, its paint worn away to bare metal in a hand-shaped arc around the latch. A short screwdriver rests on the skirting ledge underneath.
Four months unattended. Twice a week, someone quietly decides that word still applies. More in the hall of fame

Drawn by an image model.

Corrections

Nothing to correct.

Related issues

Subscribe

Get the brief in your inbox

Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.

  • Weekdays at 8:45am IST, one lead story and 6 to 9 items.
  • Sundays, an argued synthesis rather than a recap.
  • One click to leave, and quiet days say so in the subject line.
How often

Weekdays 8:45am IST + Sundays. Unsubscribe in one click.

You're asking for The Agentic Brief by email at the cadence you picked. You can unsubscribe in one click from any issue, and your address is never sold or shared.