Computer Science editorial
PatchBench: Evaluating AI Agents for Vulnerability Patching
The core problem
Innovation
Across 11 state-of-the-art agents, including the top three AIxCC agents, the original PoC-only validation inflates the patching task solve rate of agents by 1.83 on average. This means that when patches are validated only by checking whether the PoC still triggers a crash, agents appear to solve 83% more tasks than they actually do when root-cause correctness is required. The results reveal key limitations of current patching agents. Specifically:
- 25% of agent patches show substantial similarity to historical developer patches, indicating memorization.
- Agents frequently patch on the crash stack trace to suppress the crash, rather than fixing the root cause.
- PoC-only validation inflates solve rates by 1.83 on average.
These findings hold across a diverse set of agents, including those from the AIxCC competition, suggesting that the issue is systemic. The benchmark demonstrates that current evaluation practices are insufficient for measuring true vulnerability repair capabilities.
Why it matters
Who should read this
Opening member contentโฆ