Jadwal Sholat

Memuat jadwal sholat…

Computer Science editorial

Open AccessOA2026

SemRepo: A Knowledge Graph for Research Software and Its Scholarly Ecosystem

An RDF knowledge graph with 81 million triples linking 200,000 GitHub repositories to scholarly knowledge graphs for reproducibility and sustainability analysis
Abdul Rafay; Yuni Susanti; David Lamprecht; Michael Färber· 2026· DOI 10.48550/arXiv.2605.13310

The core problem

Research software is a cornerstone of modern science, yet its connection to the scholarly record remains fragmented. Repository-level metadata—such as contributors, issues, and programming languages—typically resides on platforms like GitHub, while publication metadata, author profiles, and research artifacts are scattered across separate scholarly knowledge graphs. This fragmentation hinders comprehensive analyses of software provenance, reproducibility, and sustainability.

To address this gap, the authors introduce SemRepo, an RDF knowledge graph that unifies research software with its scholarly context. SemRepo comprises over 81 million triples describing nearly 200,000 GitHub repositories associated with scientific research. By interlinking repository information with external scholarly knowledge graphs—including SemOpenAlex for author profiles, LPWC for scholarly publications, and MLSea-KG for research artifacts—SemRepo enables queries that span publications and their scholarly artifacts. This integration supports analyses that are difficult or impossible with existing resources in isolation, such as provenance reconstruction across repositories and publications, and the syst

Innovation

SemRepo successfully integrates repository-level metadata with scholarly knowledge graphs, resulting in a comprehensive resource for analyzing research software. Key quantitative outcomes include:

- **Scale**: Over 81 million triples and nearly 200,000 GitHub repositories.
- **Interlinking**: Authors linked to SemOpenAlex profiles, repositories connected to LPWC publications, and artifacts linked via MLSea-KG.
- **Query capabilities**: The graph supports complex queries that span multiple domains, such as finding all publications associated with a repository, or identifying datasets used in experiments related to a specific software project.

The following Mermaid diagram illustrates the high-level architecture of SemRepo and its connections to external knowledge graphs:

This a

Research software is a cornerstone of modern science, yet its connection to the scholarly record remains fragmented. Repository-level metadata—such as contributors, issues, and programming languages—typically resides on platforms like GitHub, while publication metadata, author profiles, and research artifacts are scattered across separate scholarly knowledge graphs. This fragmentation hinders comprehensive analyses of software provenance, reproducibility, and sustainability.
To address this gap, the authors introduce SemRepo, an RDF knowledge graph that unifies research software with its scholarly context. SemRepo comprises over 81 million triples describing nearly 200,000 GitHub repositories associated with scientific research. By interlinking repository information with external scholarly knowledge graphs—including SemOpenAlex for author profiles, LPWC for scholarly publications, and MLSea-KG for research artifacts—SemRepo enables queries that span publications and their scholarly artifacts. This integration supports analyses that are difficult or impossible with existing resources in isolation, such as provenance reconstruction across repositories and publications, and the systematic identification of risks to research reproducibility and software sustainability.

Why it matters

SemRepo addresses a critical need in the scholarly ecosystem by providing a unified graph that connects research software with its scholarly context. The integration of repository metadata with author profiles, publications, and artifacts allows for analyses that were previously impractical due to data fragmentation.

One key application is provenance reconstruction: tracing the lineage of a software project from its initial commit through its associated publications and datasets. This can help in understanding the impact and evolution of research software. Another application is the systematic identification of risks to reproducibility and sustainability. For example, by analyzing dependencies, issue histories, and contributor activity, researchers can identify repositories that are at risk of becoming unmaintained or that lack proper documentation for reproducibility.

The graph's scale—81 million triples—demonstrates its potential for large-scale analyses. However, challenges remain, such as ensuring data quality and handling the dynamic nature of repositories. Future work could expand the graph to include more platforms and artifact types, and develop user-friendly query interfaces.

In summary, SemRepo provides an important infrastructure for the large-scale analysis of software within the broader scientific research ecosystem, enabling new insights into the interplay between software and scholarly outputs.

Who should read this

CS practitioners and researchers

Opening member content…