Jadwal Sholat

Memuat jadwal sholatโ€ฆ

Ilmu Komputer & AI editorial

Open AccessOA2026

AGENTQ: Quantization-Conditioned Backdoor Attacks on LLM Agents

First study of quantization-conditioned attacks against LLM agents reveals up to 100% attack success rate with minimal utility loss
Xiaoqun Liu; Qiben Yanยท 2026ยท DOI 10.48550/arXiv.2609.14060

The core problem

Quantization has become a default deployment path for open-weight large language model (LLM) agents, enabling efficient inference on resource-constrained hardware. However, quantization is not behavior-preserving: an adversary can release a full-precision checkpoint that passes standard audits yet exhibits malicious behavior once quantized. This phenomenon is termed a **quantization-conditioned attack (QCA)**. Prior QCA research has focused on free-text generation, where harm is mediated by a human reader who may detect and filter malicious outputs. In contrast, the agentic setting poses a more severe risk: the triggered payload is a structured function that can be executed without human oversight, potentially leading to direct system compromise or data exfiltration. This paper presents the first study of QCA against LLM agents. The authors find that directly adapting prior backdoor-injection methods can produce malicious behavior after quantization, but substantially degrades benign utility, rendering the resulting attacks impractical. To understand the true upper bound of the threat, they propose AGENTQ, an attack framework that combines layer-banded LoRA injection with partial-P

Innovation

The authors evaluate AGENTQ across three trigger-action pairs and three quantization codebooks: NF4, FP4, and INT8. The trigger-action pairs are designed to represent realistic agentic tasks where a specific input triggers a malicious function call. The results show that AGENTQ achieves up to **100% post-quantization attack success rate (ASR)** while maintaining benign utility close to that of the original model.

Specifically, for each codebook, the ASR is measured as the percentage of trigger inputs that successfully cause the agent to execute the malicious action after quantization. The benign utility is measured using standard benchmarks for agentic tasks, such as task completion rate and accuracy on benign inputs. The authors report that AGENTQ incurs minimal loss of benign utility, with the quantized model performing nearly as well as the full-precision model on benign tasks.

In contrast, directly adapting prior backdoor-injection methods (e.g., standard backdoor attacks without quantization awareness) results in either low ASR after quantization or significant degradation of benign utility. For instance, a naive backdoor might achieve high ASR in full precision but fail to

Quantization has become a default deployment path for open-weight large language model (LLM) agents, enabling efficient inference on resource-constrained hardware. However, quantization is not behavior-preserving: an adversary can release a full-precision checkpoint that passes standard audits yet exhibits malicious behavior once quantized. This phenomenon is termed a **quantization-conditioned attack (QCA)**. Prior QCA research has focused on free-text generation, where harm is mediated by a human reader who may detect and filter malicious outputs. In contrast, the agentic setting poses a more severe risk: the triggered payload is a structured function that can be executed without human oversight, potentially leading to direct system compromise or data exfiltration. This paper presents the first study of QCA against LLM agents. The authors find that directly adapting prior backdoor-injection methods can produce malicious behavior after quantization, but substantially degrades benign utility, rendering the resulting attacks impractical. To understand the true upper bound of the threat, they propose AGENTQ, an attack framework that combines layer-banded LoRA injection with partial-PGD repair over a multi-codebook quantization-equivalence class. AGENTQ preserves normal agentic capability while concentrating malicious behavior in the quantized model. Across three trigger-action pairs and three codebooks (NF4, FP4, INT8), AGENTQ reaches up to 100% post-quantization attack success rate with minimal loss of benign utility, underscoring the need to make quantization-aware safety evaluation a standard requirement before open-weight agents are deployed.
AGENTQ is designed to inject a backdoor that activates only after quantization, while maintaining high benign utility in both full-precision and quantized models. The framework consists of two key components: **layer-banded LoRA injection** and **partial-PGD repair** over a multi-codebook quantization-equivalence class.

Why it matters

The findings of this study have profound implications for the security of LLM agents. The fact that an adversary can release a full-precision checkpoint that passes standard safety audits yet becomes malicious after quantization means that current safety evaluation practices are insufficient. Quantization-aware safety evaluation must become a standard requirement before open-weight agents are deployed.

The AGENTQ framework highlights a fundamental tension between model efficiency and security. Quantization is essential for deploying large models on edge devices, but it also introduces a new attack surface. The layer-banded LoRA injection and partial-PGD repair techniques are not specific to any particular model architecture, suggesting that the threat is general and could affect a wide range of open-weight agents.

Moreover, the agentic setting amplifies the risk because the triggered payload is a structured function that can be executed without human oversight. Unlike free-text generation, where a human might notice and intervene, an agent can directly perform actions such as sending emails, modifying files, or executing code. This makes the consequences of a successful attack potentially severe.

The authors suggest several directions for future work, including developing quantization-aware defenses, such as robust quantization schemes that preserve safety properties, and designing auditing procedures that can detect quantization-conditioned backdoors. They also call for the community to adopt quantization-aware safety evaluations as a standard practice.

In conclusion, AGENTQ demonstrates that quantization-conditioned backdoor attacks on LLM agents are a real and practical threat. The ability to achieve up to 100% ASR with minimal utility loss underscores the need for immediate attention to this issue. As open-weight agents become more prevalent, ensuring their safety under quantization is critical.

Who should read this

CS practitioners and researchers

Opening member contentโ€ฆ