Computer Science editorial
Open AccessOA2026
Overlay_dx: Automating Forecasting Evaluation
A visual metric for time series prediction models that combines confidence intervals and area under the curve
Long Ngo; Mohammed Amine Chamli; Jonathan Rivalan; Thomas Jaillonยท 2026ยท DOI 10.48550/arXiv.2609.24586
The core problem
Traditional evaluation metrics for time series forecasting, such as MAE, RMSE, or MAPE, provide numerical values but often lack comprehensibility. These scalar summaries can hinder effective differentiation of model performances, especially when models have similar error magnitudes but different error distributions. The authors argue that a visual metric can complement numerical assessments by showing how predictions align with actual values over time and across confidence levels. They introduce **overlay_dx**, a metric that represents the percentage of predictions falling within a confidence interval around actual values. Once evaluation results are plotted, overlay_dx computes the area under the overlay curve, providing a quantitative measure of alignment between predicted and actual values across different thresholds and predictions. The goal is to offer a unified evaluation framework that combines visual and numerical assessments, enabling improved model comparison and insights for optimization in time series prediction.
Innovation
The authors conducted extensive experiments to validate overlay_dx. They compared multiple time series prediction models, including statistical methods (e.g., ARIMA), machine learning models (e.g., gradient boosting), and deep learning architectures (e.g., LSTM, Transformer). The results show that overlay_dx and its AUC provide a more nuanced differentiation of model performance than traditional metrics. For instance, two models with similar RMSE can have significantly different overlay curves: one may consistently cover a high percentage of predictions at small thresholds, while the other may only achieve high coverage at larger thresholds. The visual overlay plot reveals patterns such as systematic bias (predictions consistently above or below actuals) and heteroscedastic errors (varying coverage across time). The AUC values correlate with but are not identical to traditional error metrics, offering complementary information. The authors report that overlay_dx is robust to outliers and scale, as it is based on relative coverage rather than absolute errors. They also demonstrate that the metric can be used for hyperparameter tuning: models optimized for overlay_dx AUC achieve bett
Traditional evaluation metrics for time series forecasting, such as MAE, RMSE, or MAPE, provide numerical values but often lack comprehensibility. These scalar summaries can hinder effective differentiation of model performances, especially when models have similar error magnitudes but different error distributions. The authors argue that a visual metric can complement numerical assessments by showing how predictions align with actual values over time and across confidence levels. They introduce **overlay_dx**, a metric that represents the percentage of predictions falling within a confidence interval around actual values. Once evaluation results are plotted, overlay_dx computes the area under the overlay curve, providing a quantitative measure of alignment between predicted and actual values across different thresholds and predictions. The goal is to offer a unified evaluation framework that combines visual and numerical assessments, enabling improved model comparison and insights for optimization in time series prediction.
The overlay_dx metric is defined as follows. For a given time series of actual values and predictions
, a confidence interval is constructed around each actual value:
, where is a threshold parameter. A prediction is considered "covered" if
. The overlay_dx at threshold is the proportion of covered predictions:
Why it matters
The overlay_dx metric addresses a key gap in forecasting evaluation: the need for interpretable, visual, and quantitative assessment. By combining a coverage-based curve with an AUC summary, it bridges the gap between visual inspection and numerical comparison. The authors discuss several advantages: (1) **Comprehensibility**: the overlay plot is intuitive, showing exactly where and when predictions fall within acceptable bounds. (2) **Flexibility**: the threshold can be set based on domain requirements (e.g., acceptable error margins). (3) **Automation**: the entire evaluation can be automated, making it suitable for large-scale model selection. Limitations include the need to choose and the fact that the metric assumes a symmetric confidence interval around actuals. Future work includes extending overlay_dx to probabilistic forecasts and multivariate time series. The authors conclude that overlay_dx offers a unified evaluation framework that can improve model comparison and provide valuable insights for further research and optimization in time series prediction.
Who should read this
CS practitioners and researchers
Opening member contentโฆ