Ilmu Komputer & AI editorial
AGENTQ: Quantization-Conditioned Backdoor Attacks on LLM Agents
The core problem
Innovation
The authors evaluate AGENTQ across three trigger-action pairs and three quantization codebooks: NF4, FP4, and INT8. The trigger-action pairs are designed to represent realistic agentic tasks where a specific input triggers a malicious function call. The results show that AGENTQ achieves up to **100% post-quantization attack success rate (ASR)** while maintaining benign utility close to that of the original model.
Specifically, for each codebook, the ASR is measured as the percentage of trigger inputs that successfully cause the agent to execute the malicious action after quantization. The benign utility is measured using standard benchmarks for agentic tasks, such as task completion rate and accuracy on benign inputs. The authors report that AGENTQ incurs minimal loss of benign utility, with the quantized model performing nearly as well as the full-precision model on benign tasks.
In contrast, directly adapting prior backdoor-injection methods (e.g., standard backdoor attacks without quantization awareness) results in either low ASR after quantization or significant degradation of benign utility. For instance, a naive backdoor might achieve high ASR in full precision but fail to
Why it matters
The findings of this study have profound implications for the security of LLM agents. The fact that an adversary can release a full-precision checkpoint that passes standard safety audits yet becomes malicious after quantization means that current safety evaluation practices are insufficient. Quantization-aware safety evaluation must become a standard requirement before open-weight agents are deployed.
The AGENTQ framework highlights a fundamental tension between model efficiency and security. Quantization is essential for deploying large models on edge devices, but it also introduces a new attack surface. The layer-banded LoRA injection and partial-PGD repair techniques are not specific to any particular model architecture, suggesting that the threat is general and could affect a wide range of open-weight agents.
Moreover, the agentic setting amplifies the risk because the triggered payload is a structured function that can be executed without human oversight. Unlike free-text generation, where a human might notice and intervene, an agent can directly perform actions such as sending emails, modifying files, or executing code. This makes the consequences of a successful attack potentially severe.
The authors suggest several directions for future work, including developing quantization-aware defenses, such as robust quantization schemes that preserve safety properties, and designing auditing procedures that can detect quantization-conditioned backdoors. They also call for the community to adopt quantization-aware safety evaluations as a standard practice.
In conclusion, AGENTQ demonstrates that quantization-conditioned backdoor attacks on LLM agents are a real and practical threat. The ability to achieve up to 100% ASR with minimal utility loss underscores the need for immediate attention to this issue. As open-weight agents become more prevalent, ensuring their safety under quantization is critical.
Who should read this
Opening member contentโฆ