Ilmu Komputer & AI editorial
Open AccessOA2026
Pick Your Poison: Learning to Select Poison Sets for Stronger LLM Backdoor Attacks
SAILS: Set-level Audit-Informed Iterative Learned Selection for Oracle-Budgeted Poison Set Optimization
Aashiq Muhamed; Mona T. Diab; Virginia Smith; Andrew Ilyas; Matthew Jagielskiยท 2026ยท DOI 10.48550/arXiv.2609.15029
The core problem
Backdoor poisoning attacks inject poisoned examples into otherwise-clean finetuning data, pairing a trigger with a target behavior that the model learns to produce when the trigger appears. Existing evaluations typically fix the number of poisoned examples and sample them at random from a candidate pool. The authors show that this practice can severely underestimate worst-case vulnerability: across three LLaMA-3-8B backdoor settings, holding the model, clean data, and poison count fixed, attack success ranges from 3% to 80% depending only on which poison set is chosen. This motivates formalizing poison selection as an optimization problem under a limited oracle budget. The paper introduces SAILS (Set-level Audit-Informed Iterative Learned Selection), which learns a set scorer from a few hundred finetune-and-evaluate runs, ranks millions of candidate sets, and audits only a small shortlist. SAILS improves held-out attack success by 30 percentage points on average over the strongest influence baselines, transfers from small-scale to full-scale finetuning, and extends to code-generation, agentic, and API-only backdoors.
Innovation
Across three LLaMA-3-8B backdoor settings, holding the model, clean data, and poison count fixed, attack success ranges from 3% to 80% depending only on which poison set is chosen. This demonstrates that random sampling severely underestimates worst-case vulnerability. SAILS improves held-out attack success by 30 percentage points on average over the strongest influence baselines. The method transfers from small-scale to full-scale finetuning, meaning that a scorer learned on a smaller model or dataset can guide selection for a larger one. It also extends to code-generation, agentic, and API-only backdoors, showing broad applicability. The oracle budget is only a few hundred finetune-and-evaluate runs, making the approach practical. The results highlight that poison selection is a critical but previously overlooked dimension of backdoor attack strength.
Backdoor poisoning attacks inject poisoned examples into otherwise-clean finetuning data, pairing a trigger with a target behavior that the model learns to produce when the trigger appears. Existing evaluations typically fix the number of poisoned examples and sample them at random from a candidate pool. The authors show that this practice can severely underestimate worst-case vulnerability: across three LLaMA-3-8B backdoor settings, holding the model, clean data, and poison count fixed, attack success ranges from 3% to 80% depending only on which poison set is chosen. This motivates formalizing poison selection as an optimization problem under a limited oracle budget. The paper introduces SAILS (Set-level Audit-Informed Iterative Learned Selection), which learns a set scorer from a few hundred finetune-and-evaluate runs, ranks millions of candidate sets, and audits only a small shortlist. SAILS improves held-out attack success by 30 percentage points on average over the strongest influence baselines, transfers from small-scale to full-scale finetuning, and extends to code-generation, agentic, and API-only backdoors.
The authors formalize poison selection as oracle-budgeted set optimization. Let
be the clean finetuning dataset,
a poison set drawn from a candidate pool
, and
the attack success rate after finetuning on
. The goal is to solve
Why it matters
The paper's central insight is that poison selection is a set-level optimization problem, not a random sampling task. By learning a set scorer from a few hundred oracle evaluations, SAILS can rank millions of candidate sets and audit only a small shortlist. This reduces the oracle budget while improving attack success. The 30 percentage point average improvement over influence baselines shows that influence functions, which are typically used for poison selection, are suboptimal for this task. The transfer from small-scale to full-scale finetuning suggests that the learned scorer captures generalizable properties of effective poison sets. The extension to code-generation, agentic, and API-only backdoors indicates that the vulnerability is not limited to text classification. The authors' formalization as oracle-budgeted set optimization provides a framework for future work on worst-case evaluation of backdoor defenses. The findings imply that defenses should be evaluated against adversarially selected poison sets, not random ones, to avoid overestimating robustness.
Who should read this
CS practitioners and researchers
Opening member contentโฆ