Jadwal Sholat

Memuat jadwal sholatโ€ฆ

Ilmu Komputer & AI editorial

Open AccessOA2026

CESBench: Benchmarking Large Language Models on Cryptographic Engineering Security for IoT Devices

A 380-item expert benchmark reveals that LLMs excel at recall and code but struggle to justify security verdicts in IoT cryptographic engineering.
Wenquan Zhou; An Wang; Jing Liang; Peien Feng; Jingqi Zhang; Yaoling Ding; Liehuang Zhuยท 2026ยท DOI 10.48550/arXiv.2609.21344

The core problem

Internet of Things (IoT) devices face a critical security challenge: a secure algorithm alone is insufficient because attackers with physical access can directly exploit implementation flaws, which are often impossible to patch after deployment. Large language models (LLMs) are increasingly used to build and analyze such implementations, yet existing benchmarks for cryptography and general cybersecurity do not cover cryptographic engineering. This paper introduces CESBench, a benchmark of 380 expert-written items spanning six sub-domains of cryptographic engineering security for IoT devices: side-channel, fault injection, implementation, countermeasures, evaluation, and integration. The benchmark aims to assess LLMs' competence across four task types: multiple-choice (209 items) for recall, judgment (67 items) requiring a security verdict and justification, scenario (63 items) for engineering diagnosis, and code (41 tasks) graded by 572 test cases. The authors validate CESBench by evaluating 11 open-weight and proprietary LLMs, using automatic scoring for multiple-choice and code, and an LLM judge for judgment and scenario responses, cross-checked with a second judge and human re-s

Innovation

The evaluation of 11 LLMs on CESBench reveals a wide range of performance. Composite scores span from 54.4% to 83.6%, indicating significant variation in cryptographic engineering security competence. On individual task types, the strongest models achieve near-ceiling performance on multiple-choice (98.6%) and code (95.1%), and 88.4% on scenario diagnosis. However, judgment tasks remain challenging, with the top score only 58.8%. Across all models, 88.5% of security verdicts are correct, but the justifications for these verdicts earn only 53.4% of the rubric marks. This discrepancy highlights that while models can often identify the correct security outcome, they struggle to provide adequate reasoning. The results suggest that multiple-choice and code tasks are largely solved by top models, whereas justifying a security verdict is the weakest competence. The benchmark's public release of prompts and per-item results facilitates further analysis and improvement.
Internet of Things (IoT) devices face a critical security challenge: a secure algorithm alone is insufficient because attackers with physical access can directly exploit implementation flaws, which are often impossible to patch after deployment. Large language models (LLMs) are increasingly used to build and analyze such implementations, yet existing benchmarks for cryptography and general cybersecurity do not cover cryptographic engineering. This paper introduces CESBench, a benchmark of 380 expert-written items spanning six sub-domains of cryptographic engineering security for IoT devices: side-channel, fault injection, implementation, countermeasures, evaluation, and integration. The benchmark aims to assess LLMs' competence across four task types: multiple-choice (209 items) for recall, judgment (67 items) requiring a security verdict and justification, scenario (63 items) for engineering diagnosis, and code (41 tasks) graded by 572 test cases. The authors validate CESBench by evaluating 11 open-weight and proprietary LLMs, using automatic scoring for multiple-choice and code, and an LLM judge for judgment and scenario responses, cross-checked with a second judge and human re-scoring. The study addresses a gap in LLM evaluation for cryptographic engineering, providing a public benchmark to drive progress in this critical area.
CESBench comprises 380 expert-written items designed to cover the breadth of cryptographic engineering security for IoT devices. The items are organized into six sub-domains: side-channel, fault injection, implementation, countermeasures, evaluation, and integration. Four task types target different competences: multiple-choice (209 items) tests recall; judgment (67 items) requires a security verdict and its justification; scenario (63 items) requires an engineering diagnosis; and code (41 tasks) are graded by 572 test cases. To validate the benchmark, 11 open-weight and proprietary LLMs answer every item. Multiple-choice and code responses are scored automatically, while judgment and scenario responses are evaluated by an LLM judge. The LLM judge's scores are checked against a second judge from another model family and human re-scoring to ensure reliability. The evaluation yields composite scores ranging from 54.4% to 83.6%. The top score on each task type is 98.6% for multiple choice, 95.1% for code, and 88.4% for scenario diagnosis, but only 58.8% for judgment. Across models, 88.5% of verdicts are correct, yet their justifications earn only 53.4% of the rubric marks. The benchmark, prompts, and per-item results are public.

Why it matters

The findings from CESBench underscore a critical gap in LLM capabilities for cryptographic engineering security. While models excel at recall-based tasks (multiple-choice) and code generation, their ability to justify security verdicts lags significantly. This is concerning because in real-world IoT security, understanding why a particular implementation is vulnerable is as important as identifying the vulnerability itself. The high accuracy on verdicts (88.5%) but low justification scores (53.4%) suggests that models may be relying on pattern recognition rather than deep reasoning. The benchmark's coverage of six sub-domains ensures a comprehensive assessment, but the results indicate that current LLMs are not yet reliable for autonomous security analysis in this domain. The authors emphasize the need for improved reasoning and explanation capabilities. The public availability of CESBench, including prompts and per-item results, provides a foundation for future research. The benchmark's design, with its mix of task types and expert-written items, sets a new standard for evaluating LLMs in cryptographic engineering. As IoT devices proliferate, ensuring their security through robust LLM-assisted analysis becomes increasingly vital.

Who should read this

CS practitioners and researchers

Opening member contentโ€ฆ