Computer Science editorial
Open AccessOA2026
Attack Success Rate Is Not a Number: On Measurement Validity in Agentic AI Security Evaluation
A meta-analysis of 259 agentic-security papers and an analytical power study show that cross-paper ASR comparison is currently unsupported, and propose a ten-item reporting checklist.
Chetan Pathade; Prathamesh Pawar; Shubham Patilยท 2026ยท DOI 10.48550/arXiv.2609.25173
The core problem
Attack success rate (ASR) is the headline metric in nearly every published evaluation of attacks on, and defenses for, LLM agents. The authors argue that ASR as currently used is not a single quantity but a family of metrics parameterized by six design choices that papers seldom specify and never hold constant across the literature. This measurement-validity problem matters because the field routinely compares ASR values across papers, ranks defenses, and draws security conclusions from numbers that may not be commensurable. The paper's aim is not to dispute any individual result but to supply the shared measurement contract the field has so far done without. The authors support their argument with two studies that require no proprietary access: a full-text meta-analysis of 259 agentic-security papers posted to arXiv between February 2025 and September 2026, and an analytical study of the statistical consequences of the reporting gaps they document.
Innovation
The meta-analysis finds that most papers report neither a variance estimate nor repeated runs for their headline attack metric: 58% (95% CI 44-71) in the hand-coded random sample of 50, and 65.3% by automated coding of all 259. Only 30.9% disclose enough about decoding to establish whether their evaluation was even stochastic. Of the 64 papers confirmed to use an LLM judge, 29.7% report any agreement check against human labels. The analytical study shows these omissions are not cosmetic: on a 100-instance benchmark, the minimum difference in ASR detectable at conventional power is 18.2 percentage points. Furthermore, two defenses whose true ASRs differ by 5 points are ranked in the wrong order by a single-run evaluation roughly 21% of the time. Because several of the six axes shift ASR in a system-dependent way, the resulting incomparability is not a constant offset that cancels in comparison.
Attack success rate (ASR) is the headline metric in nearly every published evaluation of attacks on, and defenses for, LLM agents. The authors argue that ASR as currently used is not a single quantity but a family of metrics parameterized by six design choices that papers seldom specify and never hold constant across the literature. This measurement-validity problem matters because the field routinely compares ASR values across papers, ranks defenses, and draws security conclusions from numbers that may not be commensurable. The paper's aim is not to dispute any individual result but to supply the shared measurement contract the field has so far done without. The authors support their argument with two studies that require no proprietary access: a full-text meta-analysis of 259 agentic-security papers posted to arXiv between February 2025 and September 2026, and an analytical study of the statistical consequences of the reporting gaps they document.
The paper combines two complementary methodologies. First, a full-text meta-analysis of 259 agentic-security papers posted to arXiv between February 2025 and September 2026. The authors hand-coded a random sample of 50 papers and also performed automated coding of all 259 papers, focusing on whether papers report variance estimates, repeated runs for their headline attack metric, decoding details sufficient to establish stochasticity, and agreement checks for LLM judges against human labels. Second, an analytical study on a 100-instance benchmark that quantifies the minimum detectable difference in ASR at conventional power and the probability that a single-run evaluation ranks two defenses incorrectly. The analysis treats ASR as a function of six design axes and examines whether the resulting incomparability behaves as a constant offset that cancels in comparison or as a system-dependent shift. No proprietary access was required for either study.
Why it matters
The authors conclude that cross-paper ASR comparison is currently unsupported. The six design axes that parameterize ASR are seldom specified and never held constant, so reported values cannot be assumed commensurable. The statistical results quantify the practical cost: an 18.2 percentage-point minimum detectable difference at conventional power means many published comparisons are underpowered, and a 21% misranking rate for a 5-point true difference means single-run evaluations frequently reverse the correct ordering of defenses. The absence of variance estimates, repeated runs, decoding details, and judge-agreement checks compounds the problem. To address each measured failure, the authors propose a ten-item reporting checklist targeted at the specific omissions they document. The checklist is intended as a shared measurement contract for agentic AI security evaluation, enabling valid comparison and replication without disputing any individual result.
Who should read this
CS practitioners and researchers
Opening member contentโฆ