Computer Science editorial
Open AccessOA2026
Ascent: An Agentic System over the Model Context Protocol for Real-World Clinical Data Analysis
An agentic architecture for epidemiological question answering over real-world clinical data, with EpiTrap benchmark and empirical gains over fixed pipelines.
Angelo Ziletti; Leonardo D'Ambrosi; Melanie Tuchardt; Tim Kondziellaยท 2026ยท DOI 10.48550/arXiv.2609.24620
The core problem
Answering epidemiological questions from real-world clinical data requires medical coding, schema-aware SQL, and validation of implicit choices about populations, denominators, and time. These tasks are challenging because clinical data are heterogeneous, often stored in native schemas that vary across institutions, and because epidemiological analyses are sensitive to subtle decisions that can introduce bias. The authors present Ascent, an agentic system designed to address these challenges by exposing medical coding, question answering, and cohort analysis through a shared Model Context Protocol (MCP) tool surface. This tool surface supports both standardized and native schemas, enabling flexible interaction with diverse data sources. To evaluate whether systems avoid recognized pharmacoepidemiological errors, the authors introduce EpiTrap, a dataset specifically constructed for this purpose. The paper compares a fixed pipeline with agents across different models and orchestrators, demonstrating that agentic approaches can substantially improve accuracy when powered by capable models. The work is motivated by real-world experience from clinical data projects, highlighting the nee
Innovation
The key quantitative result is that agentic systems, when powered by capable models, achieve an average improvement of 27 percentage points on native schemas and 20 percentage points on standardized schemas compared to a fixed pipeline. This improvement is measured on the EpiTrap dataset, which tests avoidance of pharmacoepidemiological errors. The results indicate that the agentic approach is particularly beneficial for native schemas, where heterogeneity and lack of standardization pose greater challenges. The authors also note that these accuracy gains require more tool calls and longer runtimes, suggesting that the agentic system is more resource-intensive. The comparison across models and orchestrators reveals that the choice of model and orchestrator significantly impacts performance, with capable models being essential to realize the benefits of the agentic approach. The paper does not provide specific numbers for tool calls or runtimes, but the qualitative statement indicates a trade-off. Additionally, experience from real projects highlights the system's value for feasibility assessment, diagnostic iteration, and expert-guided analysis, suggesting that beyond benchmark acc
Answering epidemiological questions from real-world clinical data requires medical coding, schema-aware SQL, and validation of implicit choices about populations, denominators, and time. These tasks are challenging because clinical data are heterogeneous, often stored in native schemas that vary across institutions, and because epidemiological analyses are sensitive to subtle decisions that can introduce bias. The authors present Ascent, an agentic system designed to address these challenges by exposing medical coding, question answering, and cohort analysis through a shared Model Context Protocol (MCP) tool surface. This tool surface supports both standardized and native schemas, enabling flexible interaction with diverse data sources. To evaluate whether systems avoid recognized pharmacoepidemiological errors, the authors introduce EpiTrap, a dataset specifically constructed for this purpose. The paper compares a fixed pipeline with agents across different models and orchestrators, demonstrating that agentic approaches can substantially improve accuracy when powered by capable models. The work is motivated by real-world experience from clinical data projects, highlighting the need for systems that support feasibility assessment, diagnostic iteration, and expert-guided analysis.
The Ascent system is built around a shared Model Context Protocol (MCP) tool surface that provides standardized access to medical coding, question answering, and cohort analysis functionalities. This tool surface abstracts away the underlying schema differences, allowing agents to interact with both native and standardized clinical data schemas. The system employs an agentic architecture where large language models (LLMs) act as agents that can call tools, reason about intermediate results, and iteratively refine their approach. The authors compare this agentic approach against a fixed pipeline baseline. The fixed pipeline likely consists of a predetermined sequence of steps without dynamic decision-making. The evaluation uses EpiTrap, a dataset designed to test whether systems avoid recognized pharmacoepidemiological errors. EpiTrap includes questions that require careful handling of populations, denominators, and time, and it is used to measure accuracy across different configurations. The experiments vary the underlying model (e.g., different LLMs) and the orchestrator (the component that manages agent-tool interactions). The authors report that with capable models, agents improve accuracy over the fixed pipeline by an average of 27 and 20 percentage points on native and standardized schemas, respectively. These gains come at the cost of more tool calls and longer runtimes, indicating a trade-off between accuracy and efficiency. The methodology also includes qualitative insights from real projects, emphasizing the system's value for feasibility assessment, diagnostic iteration, and expert-guided analysis.
Why it matters
The findings demonstrate that agentic systems over a shared MCP tool surface can substantially improve accuracy in answering epidemiological questions from real-world clinical data. The improvement is especially pronounced for native schemas, which are common in practice but less standardized. This suggests that agentic approaches can help bridge the gap between messy real-world data and rigorous epidemiological analysis. However, the increased tool calls and runtimes indicate a cost that must be managed, especially in time-sensitive or resource-constrained settings. The EpiTrap dataset provides a valuable benchmark for evaluating systems on their ability to avoid recognized pharmacoepidemiological errors, and it could be used to drive further improvements. The authors' experience from real projects underscores the importance of human-in-the-loop workflows, where the system supports feasibility assessment, diagnostic iteration, and expert-guided analysis. This aligns with the broader trend of using AI to augment rather than replace human expertise in complex domains. Future work could focus on optimizing the agentic architecture to reduce tool calls and runtimes, perhaps through better orchestration or caching. Additionally, extending the tool surface to cover more types of clinical data and analyses would enhance the system's utility. The paper's contributions include the Ascent system, the EpiTrap dataset, and empirical evidence of the benefits of agentic approaches in this domain. The results are promising but also highlight the need for careful validation and oversight when deploying such systems in clinical research.
Who should read this
CS practitioners and researchers
Opening member contentโฆ