Computer Science editorial
Open AccessOA2026
Evaluation of Pipelines for Data Integration into Knowledge Graphs
A benchmark for assessing coverage, correctness, and consistency in KG integration workflows
Marvin Hofer; Erhard Rahm· 2026· DOI 10.48550/arXiv.2605.22304
The core problem
Knowledge graphs (KGs) are increasingly used to integrate data from diverse sources, but the process of ingesting new data into an existing KG typically involves complex workflows or pipelines. Many possible pipelines exist for a given integration problem, yet there is no general approach to evaluate their overall quality and performance, making it difficult to determine the best choices. This paper addresses that gap by proposing KGI-Bench, a benchmark designed to evaluate integration pipelines that ingest different kinds of input data into an existing KG. The authors motivate the need for such a benchmark by highlighting the lack of standardized evaluation methods, which hinders reproducibility and informed decision-making in KG integration projects.
Innovation
To demonstrate the applicability and usefulness of KGI-Bench, the authors comparatively evaluate 12 pipelines. These pipelines vary in design choices, such as the order of integration steps, the handling of different input formats, and the use of specific mapping or fusion techniques. The evaluation analyzes their behavior across different input data formats and design choices. Results show that pipeline performance varies significantly depending on these factors, with trade-offs between coverage, correctness, and consistency. For example, some pipelines achieve high coverage but lower correctness, while others prioritize consistency at the expense of coverage. The benchmark successfully identifies strengths and weaknesses of each pipeline, providing insights for selecting or designing better integration workflows.
Knowledge graphs (KGs) are increasingly used to integrate data from diverse sources, but the process of ingesting new data into an existing KG typically involves complex workflows or pipelines. Many possible pipelines exist for a given integration problem, yet there is no general approach to evaluate their overall quality and performance, making it difficult to determine the best choices. This paper addresses that gap by proposing KGI-Bench, a benchmark designed to evaluate integration pipelines that ingest different kinds of input data into an existing KG. The authors motivate the need for such a benchmark by highlighting the lack of standardized evaluation methods, which hinders reproducibility and informed decision-making in KG integration projects.
The proposed benchmark, KGI-Bench, evaluates pipelines by analyzing their output—the updated KG—using three complementary quality metrics: coverage, correctness, and consistency. Coverage measures the extent to which the input data is represented in the updated KG. Correctness assesses the accuracy of the integrated data against a reference KG. Consistency checks for logical contradictions or violations of constraints within the KG. To support evaluation, the authors provide benchmark datasets for the movie domain, including a seed KG, overlapping input data in three formats (e.g., CSV, JSON, RDF), and a reference KG serving as ground truth. The benchmark is designed to be general and applicable to various integration scenarios. The evaluation process can be summarized as:
Why it matters
The paper discusses the implications of the findings for KG integration practice. The three metrics provide a multi-faceted view of pipeline quality, enabling practitioners to make informed decisions based on their specific requirements. The benchmark datasets and evaluation methodology are released to facilitate reproducibility and further research. The authors also outline potential extensions, such as incorporating additional metrics (e.g., performance, scalability) and applying the benchmark to other domains beyond movies. A key contribution is the demonstration that no single pipeline dominates across all metrics, underscoring the need for careful evaluation. The work lays a foundation for standardized benchmarking in KG integration, which can drive improvements in pipeline design and tooling. Future work may include automating pipeline selection and optimizing integration workflows using the benchmark results.
Who should read this
CS practitioners and researchers
Opening member content…