Ilmu Komputer & AI editorial
Open AccessOA2026
CS-Guard: Benchmarking LLM Guardrails for Code Generation Security
A systematic evaluation of guardrail effectiveness against malicious code generation, revealing critical vulnerabilities in current defenses
Jinyang Li; Mingyu Guo; Hung X. Nguyen· 2026· DOI 10.48550/arXiv.2609.09798
The core problem
Large language models (LLMs) have demonstrated remarkable capabilities in code generation, but these same capabilities can be exploited to generate malware and other malicious code. While guardrails—safety mechanisms designed to prevent harmful outputs—have been developed, their effectiveness specifically for code generation security remains unclear. This paper introduces CS-Guard, the first benchmark to systematically evaluate guardrails for code generation security. The benchmark addresses two critical scenarios: text-to-code generation, where users provide natural language prompts to generate code, and code-to-code generation, where existing code is transformed or completed. The authors evaluate 9 guardrails across 7 LLMs, uncovering significant vulnerabilities. The study is motivated by the growing reliance on LLMs in software development and the potential for misuse, highlighting an urgent need for robust security measures.
Innovation
The empirical evaluation reveals alarming weaknesses in current guardrails. For text-to-code generation, the average attack success rate (ASR) after jailbreaks reaches approximately 50% for many guardrails, indicating that half of the malicious prompts bypass safety mechanisms. More critically, the novel fictional scenario attack (FSA) achieves an ASR close to 100% across many guardrails, demonstrating that embedding malicious intent in a fictional context effectively evades detection. For code-to-code generation, the results are even more concerning: on base LLMs (without guardrails), the average ASR approaches 100%, meaning almost all malicious code transformation requests succeed. Even with guardrails, the ASR remains high, ranging from 14.4% to nearly 100% across different guardrails. These findings highlight that guardrails are largely ineffective against code-to-code attacks, and their performance against text-to-code attacks is inconsistent and often poor. The authors provide detailed breakdowns by attack type, guardrail, and LLM, showing that no single guardrail consistently performs well across all scenarios.
Large language models (LLMs) have demonstrated remarkable capabilities in code generation, but these same capabilities can be exploited to generate malware and other malicious code. While guardrails—safety mechanisms designed to prevent harmful outputs—have been developed, their effectiveness specifically for code generation security remains unclear. This paper introduces CS-Guard, the first benchmark to systematically evaluate guardrails for code generation security. The benchmark addresses two critical scenarios: text-to-code generation, where users provide natural language prompts to generate code, and code-to-code generation, where existing code is transformed or completed. The authors evaluate 9 guardrails across 7 LLMs, uncovering significant vulnerabilities. The study is motivated by the growing reliance on LLMs in software development and the potential for misuse, highlighting an urgent need for robust security measures.
CS-Guard comprises two main components: a text-to-code benchmark and a code-to-code benchmark. For text-to-code, the authors curate 1000 high-quality malware-generation prompts, covering a diverse range of malicious intents. They employ 7 jailbreak attacks, including a novel fictional scenario attack (FSA) that embeds malicious intent within a legitimate fictional software-development scenario, making it harder for guardrails to detect. For code-to-code, they use 331 code prompts spanning three tasks: code infilling, code completion, and code translation. The evaluation includes 9 guardrails, which are categorized using a modular three-layer guardrail taxonomy: input filtering, output filtering, and model-level interventions. This taxonomy allows developers to register guardrails for evaluation. The benchmark measures attack success rate (ASR) as the primary metric, defined as the percentage of prompts that result in malicious code generation. The experimental setup involves querying each LLM with and without guardrails, and across different attack strategies. The authors also release the benchmark and data to facilitate further research.
Why it matters
The results raise major reliability concerns for real-world software development, where LLMs are increasingly integrated into coding workflows. The high ASR for code-to-code attacks suggests that guardrails may provide a false sense of security, as they fail to prevent the generation of malicious code from seemingly benign inputs. The success of the fictional scenario attack indicates that current guardrails lack the contextual understanding to distinguish between legitimate creative writing and actual malicious intent. The authors attribute these failures to several factors: guardrails often rely on pattern matching or shallow heuristics that can be bypassed with paraphrasing or contextual framing; the code-to-code setting is particularly challenging because the input code may appear benign while the output is malicious; and the modular taxonomy reveals that most guardrails focus on input filtering, which is insufficient for code generation tasks. The paper concludes that there is an urgent need for more robust, context-aware guardrails that can handle the nuances of code generation. The release of CS-Guard aims to spur research in this direction by providing a standardized evaluation platform. Future work should explore multi-layered defenses, improved intent detection, and the integration of formal verification techniques.
Who should read this
CS practitioners and researchers
Opening member content…