Computer Science editorial
Open AccessOA2026
Graphical Causal Reasoning for Root Cause Analysis in Cloud Networks
A data-driven causal discovery approach for interpretable, time-aware root cause scoring in large-scale cloud networks
Fabien Chraim; Dominik Janzing; John Evansยท 2026ยท DOI 10.48550/arXiv.2606.13532
The core problem
Cloud-computing relies on large-scale networks which are inherently complex systems. Root cause analysis (RCA) of cloud network incidents is challenging due to the high dimensionality, dynamic nature, and intricate dependencies among network components. Traditional rule-based automation approaches struggle to keep pace with the evolving topology and diverse failure modes. This paper introduces a novel approach leveraging graph-based causal discovery techniques to address these limitations. The authors propose a spatiotemporal grouping strategy and an automation ontology to reduce the dimensionality of the problem, enabling the construction of a causal graph from binary time series data. The goal is to provide interpretable, time-aware root cause scoring that can assist network engineers in diagnosing incidents efficiently.
Innovation
The system was evaluated using a labeled dataset of 35 production incidents from a major cloud provider. The model successfully recalled the correct root cause in 85.7% of incidents and produced an exact match in 74.3%. These metrics indicate strong performance in identifying the true root cause among potential candidates. Additionally, the deployed system has been used in over 800 real-world incidents, with positive qualitative feedback from network engineers. The high recall rate suggests the method is effective at narrowing down the root cause, while the exact match rate demonstrates its precision in pinpointing the exact component or event responsible.
Cloud-computing relies on large-scale networks which are inherently complex systems. Root cause analysis (RCA) of cloud network incidents is challenging due to the high dimensionality, dynamic nature, and intricate dependencies among network components. Traditional rule-based automation approaches struggle to keep pace with the evolving topology and diverse failure modes. This paper introduces a novel approach leveraging graph-based causal discovery techniques to address these limitations. The authors propose a spatiotemporal grouping strategy and an automation ontology to reduce the dimensionality of the problem, enabling the construction of a causal graph from binary time series data. The goal is to provide interpretable, time-aware root cause scoring that can assist network engineers in diagnosing incidents efficiently.
The proposed method consists of several key steps:
Why it matters
The results highlight the practicality of a data-driven, causal approach to RCA in dynamic and large-scale operational environments. By leveraging causal discovery, the method overcomes the limitations of rule-based automation, which often fails to adapt to changing network conditions and novel failure modes. The spatiotemporal grouping and automation ontology effectively reduce dimensionality, making the problem tractable. The probabilistic inference with time lags provides interpretable scores, which is crucial for building trust with network engineers. The positive feedback from over 800 real-world incidents underscores the system's usefulness in production. However, the evaluation on 35 incidents, while promising, may not capture the full diversity of cloud network failures. Future work could involve expanding the dataset, incorporating more sophisticated causal discovery algorithms, and handling non-stationary time series. Overall, the approach represents a significant step towards automated, interpretable RCA in complex cloud networks.
Who should read this
CS practitioners and researchers
Opening member contentโฆ