Jadwal Sholat

Memuat jadwal sholatโ€ฆ

Ilmu Komputer & AI editorial

Open AccessOA2026

PIDS-Bench: Evaluating Prompt-Injection Detectors Under Over-Defense, Obfuscation, and Distribution Shift

A frozen multi-axis benchmark reveals that aggregate conceals provenance-sensitive over-defense in prompt-injection detectors
Yusuf Khalid Shire; Sang-Chul Kimยท 2026ยท DOI 10.48550/arXiv.2609.15017

The core problem

Prompt-injection detectors are commonly assessed using aggregate on in-distribution test data. This practice offers limited insight into detector behavior under distribution shift, especially on the benign side of the decision boundary, where false positives impose direct operational costs but are rarely measured. The paper introduces PIDS-Bench, a frozen multi-axis benchmark designed to jointly evaluate attack detection and benign false-positive behavior at fixed thresholds. The benchmark spans four evaluation axes: in-distribution inputs, hard-benign prompts that mimic injection structure without malicious intent, obfuscated attacks, and domain and structural distribution shifts. The authors evaluate seven detectors, including learned baselines, external prompt-injection classifiers, and broad-safety comparators, alongside a rule-based lower-bound reference. The central research question is whether high aggregate translates into robust performance across these axes, particularly regarding false positives on benign inputs.

Innovation

Multi-axis evaluation exposes a failure mode that aggregate conceals. A detector exceeding on the held-out split still misclassifies roughly one-third of an externally-sourced benign subset drawn from public corpora and restricted to security-adjacent content. Across a full threshold sweep and five training seeds, no internal detector reaches an operating point satisfying and hard-benign together on this stress distribution. Decomposing by provenance reveals that hard-negative augmentation nearly eliminates over-defense on curated stress inputs but leaves it substantially intact on externally-sourced prompts. This pattern is termed provenance-sensitive over-defense. The asymmetry holds across both fine-tuned architectures and does not diminish as the augmentation pool grows, with the externally-sourced FPR remaining far above the 0.10 target. Whether augmentation matched to the externally-sourced distribution would close this gap is untested; threshold calibration and curated-style augmentation alone do not.
Prompt-injection detectors are commonly assessed using aggregate on in-distribution test data. This practice offers limited insight into detector behavior under distribution shift, especially on the benign side of the decision boundary, where false positives impose direct operational costs but are rarely measured. The paper introduces PIDS-Bench, a frozen multi-axis benchmark designed to jointly evaluate attack detection and benign false-positive behavior at fixed thresholds. The benchmark spans four evaluation axes: in-distribution inputs, hard-benign prompts that mimic injection structure without malicious intent, obfuscated attacks, and domain and structural distribution shifts. The authors evaluate seven detectors, including learned baselines, external prompt-injection classifiers, and broad-safety comparators, alongside a rule-based lower-bound reference. The central research question is whether high aggregate translates into robust performance across these axes, particularly regarding false positives on benign inputs.
PIDS-Bench is a frozen benchmark, meaning its evaluation data and thresholds are fixed to enable reproducible comparisons. The benchmark evaluates detectors at fixed thresholds and also performs a full threshold sweep across five training seeds. The evaluation axes are:

Why it matters

The results highlight a critical gap in current evaluation practices for prompt-injection detectors. Aggregate on in-distribution data can be misleadingly high while operational false positives on benign, security-adjacent prompts remain unacceptable. The provenance-sensitive over-defense pattern suggests that detectors learn superficial cues from curated hard negatives that do not generalize to externally-sourced benign prompts. This has direct implications for deployment: in security-sensitive environments, a high false-positive rate on benign prompts can erode trust and cause alert fatigue. The authors note that threshold calibration and curated-style augmentation alone do not resolve the issue. Future work should explore augmentation strategies matched to the externally-sourced distribution, as well as more robust training objectives that explicitly penalize false positives on benign inputs across diverse provenances. The benchmark provides a foundation for such research by offering a multi-axis evaluation framework that measures both attack detection and benign false-positive behavior.

Who should read this

CS practitioners and researchers

Opening member contentโ€ฆ