Ilmu Komputer & AI editorial
Temporal Generalization and Explanation Stability of Control Flow Graph Neural Networks for Malware Detection
The core problem
Innovation
The choice of message-passing operator significantly changes robustness to the temporal shift. Every pairwise gap that survives correction separates an aggregating architecture from one built around a learned attentional readout. The ranking of models reverses: the flat control, which sees node features but no topology, is the best in-distribution model but among the worst across the temporal boundary. This indicates that a conventional benchmark would have rejected message passing. Neither recalibration nor ensembling substitutes for the operator choice. Attributions do not shift significantly, but explanation validity is architecture-specific. The most accurate operator on the later corpus is the hardest to explain. An architecture derived from these findings matches the best searched operator without the need for search. The shift affects both malware and benign classes alike, so these results are about robustness to distribution shift, not malware evolution.
Key quantitative outcomes:
- Training set: 459 CFGs from 2024-2025.
- Evaluation set: 223 CFGs from 2026.
- 12 GNN variants + 1 flat control.
- Flat control: best in-distribution, worst out-of-distribution.
- Attentional r
Why it matters
The findings highlight that temporal generalization is a critical yet often overlooked aspect of malware detection models. The reversal in ranking between in-distribution and out-of-distribution performance suggests that standard random splits can be misleading. The flat control's superior in-distribution performance but poor temporal robustness implies that topology-aware models are necessary for future-proof detection, but only if the right message-passing operator is chosen. The study shows that aggregating architectures (e.g., sum, mean, max) are less robust than those with learned attentional readouts. This has implications for model design: attention mechanisms may capture more generalizable patterns. However, the trade-off is that the most accurate operator on the later corpus is also the hardest to explain, raising concerns about interpretability in security-critical applications. Recalibration and ensembling do not mitigate the shift, emphasizing the need for careful operator selection. The derived architecture that matches the best searched operator without search demonstrates that principled design can outperform exhaustive search. Finally, the shift affects both malware and benign classes, indicating that the challenge is distribution shift in general, not malware evolution.
A Mermaid diagram illustrating the temporal evaluation pipeline:
Memuat diagramโฆ
Who should read this
Opening member contentโฆ