Jadwal Sholat

Memuat jadwal sholatโ€ฆ

Ilmu Komputer & AI editorial

Open AccessOA2026

Temporal Generalization and Explanation Stability of Control Flow Graph Neural Networks for Malware Detection

A strict temporal split reveals that message-passing operator choice, not recalibration or ensembling, determines robustness to distribution shift in CFG-based malware detectors.
Md. Asif Sajeed; Md. Nazrul Islam Mondal; Md Ashraful Hossen Akashยท 2026ยท DOI 10.48550/arXiv.2609.24280

The core problem

Malware detection is a critical task in cybersecurity, and graph neural networks (GNNs) over control flow graphs (CFGs) have shown promising results. However, detectors are typically evaluated on a random split of a corpus collected over a single period, which cannot reveal how well a model generalizes to later samples. This study addresses that limitation by employing a strict temporal split: every model is trained on one period and scored once on a later one. Two corpora of CFGs, each node carrying 37 features, were extracted statically from 1,989 Windows portable executables: 459 graphs from 2024-2025 for training and 223 from 2026 for evaluation. Twelve variants and a flat-feature control were trained on the earlier corpus. The research questions center on temporal generalization, the impact of message-passing operators, and the stability and validity of explanations under distribution shift.

Innovation

The choice of message-passing operator significantly changes robustness to the temporal shift. Every pairwise gap that survives correction separates an aggregating architecture from one built around a learned attentional readout. The ranking of models reverses: the flat control, which sees node features but no topology, is the best in-distribution model but among the worst across the temporal boundary. This indicates that a conventional benchmark would have rejected message passing. Neither recalibration nor ensembling substitutes for the operator choice. Attributions do not shift significantly, but explanation validity is architecture-specific. The most accurate operator on the later corpus is the hardest to explain. An architecture derived from these findings matches the best searched operator without the need for search. The shift affects both malware and benign classes alike, so these results are about robustness to distribution shift, not malware evolution.

Key quantitative outcomes:
- Training set: 459 CFGs from 2024-2025.
- Evaluation set: 223 CFGs from 2026.
- 12 GNN variants + 1 flat control.
- Flat control: best in-distribution, worst out-of-distribution.
- Attentional r

Malware detection is a critical task in cybersecurity, and graph neural networks (GNNs) over control flow graphs (CFGs) have shown promising results. However, detectors are typically evaluated on a random split of a corpus collected over a single period, which cannot reveal how well a model generalizes to later samples. This study addresses that limitation by employing a strict temporal split: every model is trained on one period and scored once on a later one. Two corpora of CFGs, each node carrying 37 features, were extracted statically from 1,989 Windows portable executables: 459 graphs from 2024-2025 for training and 223 from 2026 for evaluation. Twelve variants and a flat-feature control were trained on the earlier corpus. The research questions center on temporal generalization, the impact of message-passing operators, and the stability and validity of explanations under distribution shift.
The study uses a strict temporal split to evaluate GNN architectures for malware detection. CFGs are extracted statically from Windows portable executables, with each node represented by 37 features. The training set consists of 459 graphs from 2024-2025, and the evaluation set comprises 223 graphs from 2026. Twelve GNN variants and a flat-feature control (which uses node features but no topology) are trained. The message-passing operators include aggregating architectures (e.g., sum, mean, max) and those with learned attentional readouts (e.g., graph attention). The performance is measured in-distribution (on a random split of the training period) and out-of-distribution (on the later period). Additionally, explanation methods are applied to assess attribution stability and validity across architectures. The experimental design allows for a direct comparison of operator choices, recalibration techniques, and ensembling strategies.

Why it matters

The findings highlight that temporal generalization is a critical yet often overlooked aspect of malware detection models. The reversal in ranking between in-distribution and out-of-distribution performance suggests that standard random splits can be misleading. The flat control's superior in-distribution performance but poor temporal robustness implies that topology-aware models are necessary for future-proof detection, but only if the right message-passing operator is chosen. The study shows that aggregating architectures (e.g., sum, mean, max) are less robust than those with learned attentional readouts. This has implications for model design: attention mechanisms may capture more generalizable patterns. However, the trade-off is that the most accurate operator on the later corpus is also the hardest to explain, raising concerns about interpretability in security-critical applications. Recalibration and ensembling do not mitigate the shift, emphasizing the need for careful operator selection. The derived architecture that matches the best searched operator without search demonstrates that principled design can outperform exhaustive search. Finally, the shift affects both malware and benign classes, indicating that the challenge is distribution shift in general, not malware evolution.

A Mermaid diagram illustrating the temporal evaluation pipeline:

Protocol visualization

Memuat diagramโ€ฆ

Who should read this

CS practitioners and researchers

Opening member contentโ€ฆ