Jadwal Sholat

Memuat jadwal sholat…

Computer Science editorial

Open AccessOA2026

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark

SRE-Bench exposes a critical gap: frontier LLM agents still fail at binary reverse engineering, solving only 31.5% of real-world-scale instances.
Jeremy Spence; Nicholas Assaderaghi; Jinhao Zhu; Nikil Ravi; Raluca Ada Popa; Guannan Wei; Yangruibo Ding; Zhuo Zhang· 2026· DOI 10.48550/arXiv.2608.11469

The core problem

AI agents are rapidly improving in cybersecurity capabilities when source code is available for analysis. However, much of the software most consequential to cybersecurity—including malware, firmware, and proprietary applications—is available only as binaries. Analyzing such software requires reverse engineering (RE): recovering program semantics before analysis can be meaningfully performed. Evaluating agentic RE poses a fundamental challenge: benchmark instances must be unseen as source code in the LLMs' training data to prevent models from taking shortcuts by recognizing them rather than truly analyzing them, while also matching the scale and anti-analysis protections of real software. Existing benchmarks do not jointly satisfy these requirements. To address this, the authors introduce SRE-Bench, the first realistic, contamination-free RE benchmark. Built entirely from scratch by RE experts over 5,000 hours, SRE-Bench comprises 19 private, real-world-scale programs averaging 16.9K lines of code. They further developed 44 in-house anti-analysis primitives, yielding 262 binary instances and 1,572 deterministically graded tasks. The central research question is: how well do frontie

Innovation

Evaluation across five frontier LLMs reveals that reverse engineering remains largely unsolved. The strongest model, GPT-5.6-sol, scores 61.4% per instance and fully solves only 31.5% of the instances. Other models perform worse, indicating a significant gap in binary analysis capabilities. The results show that strong source-code security capabilities do not yet transfer to binary analysis. The per-instance score can be interpreted as a measure of partial success, while the fully-solved rate indicates complete task completion. The distribution of scores across models suggests that current agents struggle with the complexity and anti-analysis protections present in real-world binaries. The benchmark's deterministic grading ensures reproducible and objective evaluation. These findings highlight that agentic RE is an important frontier for cybersecurity, and SRE-Bench provides a rigorous testbed to measure progress.
AI agents are rapidly improving in cybersecurity capabilities when source code is available for analysis. However, much of the software most consequential to cybersecurity—including malware, firmware, and proprietary applications—is available only as binaries. Analyzing such software requires reverse engineering (RE): recovering program semantics before analysis can be meaningfully performed. Evaluating agentic RE poses a fundamental challenge: benchmark instances must be unseen as source code in the LLMs' training data to prevent models from taking shortcuts by recognizing them rather than truly analyzing them, while also matching the scale and anti-analysis protections of real software. Existing benchmarks do not jointly satisfy these requirements. To address this, the authors introduce SRE-Bench, the first realistic, contamination-free RE benchmark. Built entirely from scratch by RE experts over 5,000 hours, SRE-Bench comprises 19 private, real-world-scale programs averaging 16.9K lines of code. They further developed 44 in-house anti-analysis primitives, yielding 262 binary instances and 1,572 deterministically graded tasks. The central research question is: how well do frontier LLM agents perform on realistic, contamination-free reverse engineering tasks?
SRE-Bench was constructed from scratch by reverse engineering experts with over 5,000 hours of effort. The benchmark includes 19 private, real-world-scale programs, each averaging 16.9K lines of code. To simulate real-world anti-analysis protections, the authors developed 44 in-house anti-analysis primitives. These primitives were applied to generate 262 binary instances, which were then used to create 1,572 deterministically graded tasks. The benchmark ensures contamination-free evaluation by using private programs not present in LLM training data. The evaluation spans five frontier LLMs: GPT-5.6-sol, Claude-Opus-5, GPT-5.5, Grok-4.5, and GLM-5.2. The tasks are designed to measure agents' ability to recover program semantics from binaries. The benchmark also includes controlled ablations to assess the impact of contamination control and realistic scale. The overall workflow can be represented as:

Why it matters

The analysis reveals that agents behave differently from human engineers. Specifically, agents are relatively insensitive to compiler optimization and static linking, which are common challenges in reverse engineering. This suggests that agents may be relying on patterns or heuristics rather than deep semantic understanding. Controlled ablations confirm that both contamination control and realistic scale are essential for meaningful evaluation. Without contamination control, models might recognize code from training data and take shortcuts. Without realistic scale, the benchmark may not reflect the challenges of real-world software. The results indicate that current agents are not yet capable of fully automating reverse engineering tasks. The authors emphasize that SRE-Bench is a crucial step toward measuring and advancing agentic cybersecurity. Future work may involve developing more sophisticated agents that can handle the complexities of binary analysis. The benchmark's design, including the use of private programs and anti-analysis primitives, sets a new standard for evaluating RE capabilities. The findings underscore the need for continued research in this area to bridge the gap between source-code and binary analysis.

Who should read this

CS practitioners and researchers

Opening member content…