Jadwal Sholat

Memuat jadwal sholat…

Ilmu Komputer & AI editorial

Open AccessOA2026

Prediction Is Not Detection: Evaluating Pre-Recognition Claims in Longitudinal Clinical AI

A methodological framework for testing whether longitudinal clinical AI detects disease before clinical recognition, not merely predicts recognition-mediated events
Jing Yang; Long R. Jiao; Xiujun Cai; Zongjiu Zhang· 2026· DOI 10.48550/arXiv.2609.25852

The core problem

Longitudinal clinical AI systems are increasingly claimed to enable early detection of disease, yet the evidentiary basis for such claims is often weaker than it appears. The central problem is that clinically useful early detection requires validated pre-recognition lead time: the interval between when a model flags risk and when a clinician would otherwise recognize the condition. Many evaluations, however, are event-based and inadvertently allow recognition-mediated care-process signals to act as shortcuts. For example, a model may learn from orders, referrals, or documentation patterns that already reflect a clinician's suspicion, rather than from pathophysiological signals that precede recognition. When recognition-dependent endpoints are used as reference standards, apparent performance and lead time can be inflated, and cross-center transportability is undermined. Such results may still serve prognosis, but they do not establish detection before recognition. The authors argue that the claim of pre-recognition detection must be made testable through explicit definitions and independent reference standards. They introduce three methodological components: an interval-censored p

Innovation

The abstract does not report empirical results from a specific dataset; instead, it presents a conceptual and methodological contribution. The authors argue that event-based evaluations can inflate apparent performance and lead time. By formalizing the pre-recognition transition, they show that many reported lead times may actually reflect recognition-mediated signals rather than true early detection. The proposed framework yields a testable claim: a model truly detects before recognition only if its flags occur before the recognition proxy and are validated against an independent as-of reference standard. When these conditions are not met, the model's output may still be useful for prognosis, but it cannot be interpreted as detection. The authors also note that recognition-dependent endpoints undermine cross-center transport, because recognition practices vary across sites. Thus, a model that appears to detect early at one center may fail at another if it relies on site-specific care-process shortcuts. The result is a clear distinction: prediction of recognition-mediated events is not equivalent to detection of pre-recognition disease. The framework provides a way to quantify this
Longitudinal clinical AI systems are increasingly claimed to enable early detection of disease, yet the evidentiary basis for such claims is often weaker than it appears. The central problem is that clinically useful early detection requires validated pre-recognition lead time: the interval between when a model flags risk and when a clinician would otherwise recognize the condition. Many evaluations, however, are event-based and inadvertently allow recognition-mediated care-process signals to act as shortcuts. For example, a model may learn from orders, referrals, or documentation patterns that already reflect a clinician's suspicion, rather than from pathophysiological signals that precede recognition. When recognition-dependent endpoints are used as reference standards, apparent performance and lead time can be inflated, and cross-center transportability is undermined. Such results may still serve prognosis, but they do not establish detection before recognition. The authors argue that the claim of pre-recognition detection must be made testable through explicit definitions and independent reference standards. They introduce three methodological components: an interval-censored pre-recognition transition, an independent as-of reference standard, and a prespecified recognition proxy. Together, these elements aim to separate genuine pre-recognition signal from recognition-mediated shortcuts, providing a rigorous framework for evaluating longitudinal clinical AI.

The authors propose a formal evaluation framework built on three pillars. First, they define an interval-censored pre-recognition transition. Let denote the time of clinical recognition and denote the time at which a pre-recognition pathophysiological transition occurs. Because is not directly observed, it is interval-censored between the last negative assessment and the first positive recognition. The target of detection is the event

, i.e., the model must flag risk before recognition. Second, they specify an independent as-of reference standard. Instead of using the eventual diagnosis as the gold standard, the reference standard is defined as-of a given time point, using only information available up to that time and independent of the model's predictions. This avoids circularity where recognition-dependent endpoints leak into the label. Third, they prespecify a recognition proxy: a measurable surrogate for the recognition process, such as the first clinical action (e.g., diagnostic test order or referral) that indicates suspicion. The proxy is fixed before analysis to prevent post hoc tuning. The evaluation then compares model predictions against the as-of reference standard, with lead time defined as
for true pre-recognition detections. The framework can be summarized as a flow:

Why it matters

The core insight of this work is that prediction is not detection. In longitudinal clinical AI, a model can achieve high accuracy by predicting events that are already recognition-mediated, such as orders or diagnoses, without ever detecting disease before a clinician would. This creates a shortcut that inflates performance and lead time, and it makes cross-center transport unreliable because recognition patterns differ. The authors' three-part framework—interval-censored pre-recognition transition, independent as-of reference standard, and prespecified recognition proxy—addresses this by making the claim of pre-recognition detection falsifiable. The interval-censored transition acknowledges that the true biological onset is unobserved, so the evaluation must reason about the interval between negative and positive assessments. The as-of reference standard prevents leakage from future recognition events. The prespecified recognition proxy ensures that the definition of recognition is not tuned to the model's outputs. Together, these elements shift the evaluation from event prediction to detection before recognition. The discussion also implies that many published claims of early detection may need re-evaluation. For clinical AI to deliver on its promise, evaluations must distinguish prognosis from detection. The framework is general and can be applied to any longitudinal clinical AI system where pre-recognition lead time is claimed. Limitations include the need for a valid recognition proxy and sufficient longitudinal data to define the interval-censored transition. Future work should apply this framework empirically across multiple centers to test transportability and to quantify the gap between prediction and detection.

Who should read this

CS practitioners and researchers

Opening member content…