Computer Science editorial
SemRepo: A Knowledge Graph for Research Software and Its Scholarly Ecosystem
The core problem
Research software is a cornerstone of modern science, yet its connection to the scholarly record remains fragmented. Repository-level metadata—such as contributors, issues, and programming languages—typically resides on platforms like GitHub, while publication metadata, author profiles, and research artifacts are scattered across separate scholarly knowledge graphs. This fragmentation hinders comprehensive analyses of software provenance, reproducibility, and sustainability.
To address this gap, the authors introduce SemRepo, an RDF knowledge graph that unifies research software with its scholarly context. SemRepo comprises over 81 million triples describing nearly 200,000 GitHub repositories associated with scientific research. By interlinking repository information with external scholarly knowledge graphs—including SemOpenAlex for author profiles, LPWC for scholarly publications, and MLSea-KG for research artifacts—SemRepo enables queries that span publications and their scholarly artifacts. This integration supports analyses that are difficult or impossible with existing resources in isolation, such as provenance reconstruction across repositories and publications, and the syst
Innovation
SemRepo successfully integrates repository-level metadata with scholarly knowledge graphs, resulting in a comprehensive resource for analyzing research software. Key quantitative outcomes include:
- **Scale**: Over 81 million triples and nearly 200,000 GitHub repositories.
- **Interlinking**: Authors linked to SemOpenAlex profiles, repositories connected to LPWC publications, and artifacts linked via MLSea-KG.
- **Query capabilities**: The graph supports complex queries that span multiple domains, such as finding all publications associated with a repository, or identifying datasets used in experiments related to a specific software project.
The following Mermaid diagram illustrates the high-level architecture of SemRepo and its connections to external knowledge graphs:
This a
Why it matters
SemRepo addresses a critical need in the scholarly ecosystem by providing a unified graph that connects research software with its scholarly context. The integration of repository metadata with author profiles, publications, and artifacts allows for analyses that were previously impractical due to data fragmentation.
One key application is provenance reconstruction: tracing the lineage of a software project from its initial commit through its associated publications and datasets. This can help in understanding the impact and evolution of research software. Another application is the systematic identification of risks to reproducibility and sustainability. For example, by analyzing dependencies, issue histories, and contributor activity, researchers can identify repositories that are at risk of becoming unmaintained or that lack proper documentation for reproducibility.
The graph's scale—81 million triples—demonstrates its potential for large-scale analyses. However, challenges remain, such as ensuring data quality and handling the dynamic nature of repositories. Future work could expand the graph to include more platforms and artifact types, and develop user-friendly query interfaces.
In summary, SemRepo provides an important infrastructure for the large-scale analysis of software within the broader scientific research ecosystem, enabling new insights into the interplay between software and scholarly outputs.
Who should read this
Opening member content…