Jadwal Sholat

Memuat jadwal sholatโ€ฆ

Ilmu Komputer & AI editorial

Open AccessOA2026

HE-Guardrail: A Homomorphic Guardrail Against Jailbreak Attacks for Encrypted Large Language Model Inference

Evaluating safety guardrails entirely over encrypted data to protect privacy-preserving LLM inference from malicious clients
Byeongseo Min; Yongwoo Lee; Young-Sik Kim; Yongjune Kimยท 2026ยท DOI 10.48550/arXiv.2609.21484

The core problem

Homomorphic encryption (HE) has emerged as a promising approach to privacy-preserving machine learning (PPML), enabling computation directly over encrypted data. In HE-based PPML, a client submits an encrypted input to the server, which evaluates models such as large language models (LLMs) without access to the underlying plaintext. This setting protects client data confidentiality, but the authors identify a critical security vulnerability: HE-LLM inference is vulnerable to malicious clients that submit adversarial prompts, such as jailbreak attacks. The same confidentiality that protects benign clients also prevents the server from inspecting incoming prompts or generated responses, making adversarial attempts difficult to detect or block and potentially allowing successful attacks to remain entirely invisible to the server. To address this vulnerability, the paper proposes HE-Guardrail, a framework that evaluates guardrail mechanisms entirely over encrypted data and homomorphically controls whether the target-model response is returned to the client. The work instantiates HE-Guardrail with three representative guardrails: Llama Guard, JBShield, and GradSafe.

Innovation

The authors evaluate HE-Guardrail with three representative guardrails: Llama Guard, JBShield, and GradSafe. The results show that HE-Guardrail closely reproduces the decisions of the corresponding plaintext guardrails in the encrypted domain. This indicates that the homomorphic evaluation of guardrail logic preserves the detection performance of the original guardrails. The paper reports distinct security-efficiency-utility trade-offs across the three instantiations. For example, some guardrails may offer stronger security guarantees at the cost of higher computational overhead, while others may be more efficient but slightly less robust. The experiments likely measure metrics such as detection accuracy, false positive/negative rates, and computational latency. The exact numerical results are not provided in the abstract, but the key finding is that encrypted guardrails can match plaintext guardrail decisions, enabling effective jailbreak detection without compromising client privacy.
Homomorphic encryption (HE) has emerged as a promising approach to privacy-preserving machine learning (PPML), enabling computation directly over encrypted data. In HE-based PPML, a client submits an encrypted input to the server, which evaluates models such as large language models (LLMs) without access to the underlying plaintext. This setting protects client data confidentiality, but the authors identify a critical security vulnerability: HE-LLM inference is vulnerable to malicious clients that submit adversarial prompts, such as jailbreak attacks. The same confidentiality that protects benign clients also prevents the server from inspecting incoming prompts or generated responses, making adversarial attempts difficult to detect or block and potentially allowing successful attacks to remain entirely invisible to the server. To address this vulnerability, the paper proposes HE-Guardrail, a framework that evaluates guardrail mechanisms entirely over encrypted data and homomorphically controls whether the target-model response is returned to the client. The work instantiates HE-Guardrail with three representative guardrails: Llama Guard, JBShield, and GradSafe.
HE-Guardrail operates in the encrypted domain, where the server holds encrypted model weights and receives encrypted client inputs. The framework integrates a guardrail model that evaluates the encrypted prompt and/or response and produces an encrypted decision signal. This signal is then used to homomorphically control whether the target-model response is returned to the client. Formally, let denote the encryption function and the decryption function. The server computes:

Why it matters

The work highlights a fundamental tension in privacy-preserving machine learning: the same encryption that protects benign clients also blinds the server to malicious inputs. HE-Guardrail resolves this by moving guardrail evaluation into the encrypted domain, allowing the server to enforce safety policies without decrypting user data. The security-efficiency-utility trade-offs suggest that no single guardrail is optimal for all scenarios; system designers must choose based on their priorities. For instance, Llama Guard may provide robust detection but require significant homomorphic computation, while JBShield or GradSafe might offer lighter-weight alternatives. The framework's ability to homomorphically control response release ensures that even if a jailbreak attempt succeeds in generating a harmful response, the server can block its delivery. This adds a layer of defense that is invisible to the client, as the client only receives either a valid response or a blocked signal. The approach is general and could be extended to other guardrails or safety mechanisms. Future work may explore optimizing the homomorphic evaluation of guardrails to reduce latency and resource consumption, as well as integrating HE-Guardrail with other PPML techniques. Overall, HE-Guardrail represents a significant step toward secure and private LLM inference, addressing a critical vulnerability in existing HE-based PPML systems.

Who should read this

CS practitioners and researchers

Opening member contentโ€ฆ