Ilmu Komputer & AI editorial
Passive Hybrid Network-Based Intrusion Detection System (Hybrid-NIDS) Combining Suricata and Random Forest
The core problem
Innovation
On the prepared hold-out set, RF-41 achieved an score of 0.971360 and a ROC-AUC of 0.999671. The NFStream-compatible RF-21 model achieved an score of 0.970148 on the same boundary. These results indicate near-perfect classification performance on the benchmark dataset, with minimal degradation from feature reduction.
However, operational validation revealed a stark contrast. On a labeled laboratory PCAP, both RF-21 and the strictly correlated branch achieved a recall of only 0.0095—meaning the system detected less than 1% of actual attacks. Furthermore, RF-21 produced no alerts across five additional 60-second attack sessions. In an unlabeled normal-traffic test, the system generated 439 alerts from 2,375 flows, yielding an alert ratio of approximately 0.185. The authors explicitly refrain from interpreting this as a false-positive rate due to the lack of ground-truth labels.
These findings underscore a severe benchmark-to-deployment domain shift. The high and ROC-AUC scores on UNSW-NB15 do not translate into operational detection capability. The recall of 0.0095 on the labeled PCAP is particularly alarming, suggesting that the model's learned features are highly specifi
Why it matters
The study's central finding is that strong performance on a public benchmark does not directly translate into operational effectiveness. The near-zero recall on the labeled PCAP and the absence of alerts in live attack sessions indicate that the Random Forest models, despite their high scores, are not detecting real intrusions in the operational environment. This domain shift likely arises from differences in traffic characteristics, feature distributions, and attack manifestations between the UNSW-NB15 dataset and the laboratory network.
The authors also note that the reported experiments do not demonstrate that Suricata-Random Forest correlation provides better operational detection than Suricata alone. This is a critical caveat: the hybrid approach, while theoretically appealing, did not yield operational benefits in this prototype. The alert ratio from normal traffic (439 alerts from 2,375 flows) suggests a high volume of alerts, but without ground-truth labels, its significance remains ambiguous.
The paper's methodological rigor in controlling for feature-duplicate leakage and separating benchmark from operational validation is commendable. It highlights the importance of domain adaptation and realistic testing in NIDS research. The current Hybrid-NIDS should be interpreted as a passive prototype and evaluation framework rather than a deployable solution. Future work should focus on bridging the domain gap, perhaps through transfer learning, adversarial validation, or incorporating more diverse training data that reflects operational conditions.
Who should read this
Opening member content…