Jadwal Sholat

Memuat jadwal sholat…

Computer Science editorial

Open AccessOA2026

Candidate Comparability Before Promotion: Conditional Validation in Adaptive Network Intrusion Detection

How challenger construction and evidence volume shape promotion decisions in adaptive NIDS
Roberto Fernández-Barrios; Iker Pastor-López; Amaia Pikatza-Huerga; Pablo García Bringas· 2026· DOI 10.48550/arXiv.2609.04388

The core problem

Adaptive network intrusion detection systems (NIDS) retrain classifiers after drift alarms, but an alarm only signals change—it does not establish that a challenger model should replace the deployed incumbent. Promotion is security-relevant because it changes the model responsible for subsequent attack detection. The authors identify a methodological problem: promotion conclusions may depend on how the challenger was constructed and on how much evidence supports it. They test this dependence across three benchmarks: CICIDS2017, UNSW-NB15, and ToN-IoT. The study uses self-contained challenger pipelines, nested candidate-size controls, a common-harness comparison of nine update policies, and a final sensitivity analysis that confines every exact feature vector to one evaluation, training, or probe role. The central question is whether promotion decisions are robust to challenger comparability and evidence volume.

Innovation

With incumbent-owned frozen preprocessing, apparent promotion harm was amplified; with self-contained challenger pipelines, the mean full-drift harm did not persist. Raising nominal candidate evidence from 512 to 2,000 samples per class improved promotion under pool-constructed progressive drift by +0.53, +1.67, and +0.38 balanced-accuracy points on CICIDS2017, UNSW-NB15, and ToN-IoT respectively. These gains were positive and statistically resolved in all three benchmarks, but materially benchmark-dependent rather than homogeneous, and driven mainly by fewer false positives. Policy conclusions were partially robust: policy ordering changed with candidate comparability, no policy globally dominated, and earlier compatibility statements for a label-free estimator and a calibrated ensemble narrowed. Validation helped evidence-disadvantaged challengers but added no average benefit at parity. Thirteen replays on real, time-ordered traffic showed no net harm from always deploying.
Adaptive network intrusion detection systems (NIDS) retrain classifiers after drift alarms, but an alarm only signals change—it does not establish that a challenger model should replace the deployed incumbent. Promotion is security-relevant because it changes the model responsible for subsequent attack detection. The authors identify a methodological problem: promotion conclusions may depend on how the challenger was constructed and on how much evidence supports it. They test this dependence across three benchmarks: CICIDS2017, UNSW-NB15, and ToN-IoT. The study uses self-contained challenger pipelines, nested candidate-size controls, a common-harness comparison of nine update policies, and a final sensitivity analysis that confines every exact feature vector to one evaluation, training, or probe role. The central question is whether promotion decisions are robust to challenger comparability and evidence volume.
The experimental design isolates two factors: challenger construction and candidate evidence size. Challenger pipelines are made self-contained to avoid incumbent-owned frozen preprocessing, which previously amplified apparent promotion harm. Candidate evidence is controlled via nested candidate-size settings, raising nominal evidence from 512 to 2,000 samples per class. Nine update policies are compared under a common harness. A final sensitivity analysis enforces strict separation of feature vectors into training, evaluation, or probe roles to prevent leakage. The benchmarks are CICIDS2017, UNSW-NB15, and ToN-IoT. The evaluation metric is balanced accuracy, with attention to false positives. The authors also replay thirteen time-ordered traffic scenarios to assess net harm from always deploying. Formally, let denote the change in balanced accuracy after promotion; the study estimates under progressive drift and full drift conditions.

Why it matters

The findings indicate that challenger construction and evidence volume are not neutral methodological choices; they directly influence promotion conclusions. Self-contained pipelines reduce artificial harm, while increased evidence benefits challengers that are otherwise disadvantaged. However, the benchmark-dependent effect sizes caution against universal promotion policies. The narrowing of compatibility statements for a label-free estimator and a calibrated ensemble suggests that earlier claims may have been overgeneralized. The absence of net harm from always deploying in time-ordered replays does not imply that promotion is always safe; rather, it highlights the need for conditional validation. The authors recommend that challenger construction and evidence be controlled, reported, and interpreted explicitly when promotion is evaluated. A flow of the promotion decision process can be represented as:

Who should read this

CS practitioners and researchers

Opening member content…