Story · Semantic Scholar
Decomposing LLM-Based Testing with Agent Skills: A Case Study on Numerical Inconsistencies (Semantic Scholar)
paper · Story page
The authors split a differential-testing pipeline into Agent Skills for program generation and feedback-guided mutation, then compare four configurations. Feedback-guided mutation gave the largest gain at 12 percentage points; adding more procedural knowledge cost 3.
In plain words
- Researchers split artificial intelligence (AI) software testing into separate instructions for creating tests and improving them.
- The tests look for programs giving different answers to the same calculation.
- Changing tests based on earlier results raised the rate of finding conflicting answers by 12 percentage points.
- For software testers, the study suggests that more detailed instructions do not always help uncover problems.
Appeared in
- VS Code rebuilt its release process around agents and now ships weekly
Oct 05, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.