Ilmu Komputer & AI editorial
Open AccessOA2026
What Makes Adversarial Examples Transfer Across Deepfake Detectors?
A controlled study of 60 detectors reveals that source–target compatibility and source-model selection are central to credible black-box robustness evaluation.
Rafael M. Mamede; Pedro C. Neto; Ana F. Sequeira· 2026· DOI 10.48550/arXiv.2609.10002
The core problem
Deepfake detectors are known to be vulnerable to transfer-based black-box attacks, where adversarial examples are crafted on a source surrogate model and then transferred to an unknown target model. However, the factors that determine whether an attack will succeed across models remain poorly understood. Prior work typically evaluates only a limited pool of detectors and rarely separates architectural factors from training-related factors. This study addresses that gap by conducting a controlled evaluation of adversarial transferability across 60 detectors, spanning six backbones, two pretraining regimes, and five training-data configurations. The authors use two attack procedures: AutoAttack (AA) and the Carlini–Wagner attack with Expectation over Transformation (CW–EOT). The central research question is: how does source–target compatibility shape attack success? The findings establish source–target compatibility and source-model selection as critical dimensions for credible transfer-based black-box robustness evaluation.
Innovation
Matched comparisons reveal significantly higher transfer when source and target share an exact backbone, architecture family, pretraining regime, or training data. The compatibility structure is attack-dependent. Under AutoAttack, exact backbone compatibility has the largest effect on transferability. Under CW–EOT, shared pretraining and training data have the largest effects. When transfer is averaged across non-target sources, the mean attack success rate (ASR) is under AA and under CW–EOT. However, a multi-source oracle that combines both attacks attains a mean ASR of after excluding exact backbone and training-data matches. This demonstrates that source averaging can substantially understate target vulnerability. The gap between single-source average and multi-source oracle is large: percentage points under CW–EOT, and percentage points under AA. These results highlight that evaluating transferability with a single source model or a small pool can severely underestimate the risk.
Deepfake detectors are known to be vulnerable to transfer-based black-box attacks, where adversarial examples are crafted on a source surrogate model and then transferred to an unknown target model. However, the factors that determine whether an attack will succeed across models remain poorly understood. Prior work typically evaluates only a limited pool of detectors and rarely separates architectural factors from training-related factors. This study addresses that gap by conducting a controlled evaluation of adversarial transferability across 60 detectors, spanning six backbones, two pretraining regimes, and five training-data configurations. The authors use two attack procedures: AutoAttack (AA) and the Carlini–Wagner attack with Expectation over Transformation (CW–EOT). The central research question is: how does source–target compatibility shape attack success? The findings establish source–target compatibility and source-model selection as critical dimensions for credible transfer-based black-box robustness evaluation.
The study evaluates adversarial transferability using a controlled experimental design. The detector pool consists of 60 models built from six backbone architectures, two pretraining regimes, and five training-data configurations. For each source–target pair, adversarial examples are generated on the source surrogate using two attack methods:
Why it matters
The findings establish source–target compatibility and source-model selection as central dimensions of credible transfer-based black-box robustness evaluation. The attack-dependent nature of compatibility effects suggests that defenders cannot rely on a single attack method to assess robustness. Under AA, architectural similarity dominates, while under CW–EOT, training-related factors dominate. This implies that robustness evaluations should include multiple attacks and diverse source models. The multi-source oracle result shows that an attacker with access to multiple surrogates and attacks can achieve high ASR even when exact backbone and training-data matches are excluded. Therefore, reporting only average ASR across non-target sources is misleading. The authors recommend that future evaluations report pairwise transfer matrices, consider source-model selection, and use multi-source oracles to estimate worst-case vulnerability. The release of 240,000 adversarial images and complete pairwise results supports reproducibility and further research. The study's controlled design disentangles architectural from training factors, providing a clear taxonomy for understanding transferability. The implications extend to deepfake detection deployment: systems should be tested against a diverse set of surrogate models and attacks to avoid overestimating robustness.
Who should read this
CS practitioners and researchers
Opening member content…