Story · arXiv (via papers.cool)
RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests (arXiv (via papers.cool))
arXiv, 4 pp · Story page
RealSWE evaluates coding agents on task variants that keep the task and gold patch fixed while varying only the request's information content and writing style.
In plain words
- Researchers tested coding systems with short, casual requests that resemble what real users write.
- They created several request versions for each task while keeping the required code change identical.
- The versions changed how much information was included and how formally the request was written.
- Realistic requests lowered average success by 6.4 percentage points and could change the ranking order.
- People choosing coding systems may get misleading comparisons from tests built around unusually detailed requests.
Appeared in
- Realistic prompts drop coding-agent scores, and tool filtering beats prompt rules
Sep 01, 2026 · lead story
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.