Jadwal Sholat

Memuat jadwal sholat…

Computer Science editorial

Open AccessOA2026

Reinforcement Learning Inspired Black-box Adversarial Attacks for Computer Vision

RIBA: A Query-Efficient Black-box Attack Using Reinforcement Learning Concepts
Florian Krone; Elena Hoemann; Sven Hallerbach· 2026· DOI 10.48550/arXiv.2609.24249

The core problem

Neural networks, including convolutional and transformer-based architectures, are fundamental to modern computer vision systems. However, they are vulnerable to adversarial perturbations—small, often imperceptible changes to input images that drastically alter model predictions. Such attacks pose a significant threat to safety-critical applications. Most existing attacks operate under the white-box threat model, assuming full access to the target model's architecture and parameters, which is unrealistic in practice. This work introduces a novel approach under the more realistic black-box threat model, where the attacker can only query the model and observe outputs. The proposed method, Reinforcement Learning Inspired Black-box Adversarial Attack (RIBA), leverages concepts from reinforcement learning to optimize perturbations against a non-differentiable target model. By exploiting the query efficiency of reinforcement learning algorithms, RIBA aims to generate successful adversarial examples with fewer queries than state-of-the-art black-box attacks.

Innovation

The authors evaluate RIBA against state-of-the-art black-box attacks on CIFAR-10 and ImageNet datasets using ResNet-18 and ViT-B/16 models, respectively. On CIFAR-10 with ResNet-18, RIBA requires 25.4% fewer median queries to generate successful adversarial examples compared to the best baseline. On ImageNet with ViT-B/16, RIBA achieves a 22.5% reduction in median queries. Furthermore, on an adversarially trained model, RIBA matches the performance of white-box attacks, demonstrating its effectiveness even against defenses. The experiments also show that RIBA maintains a high attack success rate while keeping the number of queries low, making it practical for real-world black-box scenarios. The results are summarized in the following table:

| Dataset | Model | Median Query Reduction |
|-----------|------------|------------------------|
| CIFAR-10 | ResNet-18 | 25.4% |
| ImageNet | ViT-B/16 | 22.5% |

These improvements are consistent across different perturbation budgets and model architectures, highlighting the robustness and generalizability of RIBA.

Neural networks, including convolutional and transformer-based architectures, are fundamental to modern computer vision systems. However, they are vulnerable to adversarial perturbations—small, often imperceptible changes to input images that drastically alter model predictions. Such attacks pose a significant threat to safety-critical applications. Most existing attacks operate under the white-box threat model, assuming full access to the target model's architecture and parameters, which is unrealistic in practice. This work introduces a novel approach under the more realistic black-box threat model, where the attacker can only query the model and observe outputs. The proposed method, Reinforcement Learning Inspired Black-box Adversarial Attack (RIBA), leverages concepts from reinforcement learning to optimize perturbations against a non-differentiable target model. By exploiting the query efficiency of reinforcement learning algorithms, RIBA aims to generate successful adversarial examples with fewer queries than state-of-the-art black-box attacks.
RIBA formulates the generation of adversarial perturbations as a reinforcement learning problem. The attacker interacts with the target model as an environment, receiving rewards based on the success of the attack (e.g., misclassification). The policy network is trained to output perturbations that maximize the reward while minimizing the number of queries. Since the target model is non-differentiable from the attacker's perspective, gradient-based optimization is not possible; instead, RIBA uses policy gradient methods to estimate the gradient of the expected reward with respect to the policy parameters. The overall objective can be expressed as:

Why it matters

The success of RIBA stems from its ability to learn a policy that efficiently explores the perturbation space. By framing the attack as a reinforcement learning problem, RIBA can adapt to the target model's decision boundaries without requiring gradients. The use of a replay buffer and prioritized experience replay further enhances sample efficiency, which is crucial in black-box settings where queries are expensive. The authors note that RIBA's performance on adversarially trained models is particularly promising, as it suggests that reinforcement learning can overcome some defenses that rely on gradient masking. However, RIBA's reliance on a large number of interactions during training (though not during attack) may limit its applicability in scenarios with strict query budgets. Future work could explore transferability of the learned policy across models and datasets, as well as extending RIBA to other domains such as audio or text. The following Mermaid diagram illustrates the RIBA workflow:

Overall, RIBA represents a significant step towards practical black-box adversarial attacks, balancing query efficiency with attack success.

Who should read this

CS practitioners and researchers

Opening member content…