Jadwal Sholat

Memuat jadwal sholatโ€ฆ

Computer Science editorial

Open AccessOA2026

OpenHealth Lake: Designing and testing a data lakehouse platform for health applications

A data lakehouse prototype integrating data federation and FAIR principles for collaborative global health research
Danilo Silva; Monika Moir; Cheryl Baxter; Tulio de Oliveira; Joicymara Xavier; Marcel Dunaiskiยท 2026ยท DOI 10.48550/arXiv.2605.19922

The core problem

Bioinformatics and health sciences continuously generate extensive heterogeneous datasets, making data management a complex challenge. In collaborative global health initiatives, secure storage and sharing of data are crucial to support impactful research. However, the absence of a unified data management platform complicates efficient data exchange and governance within these initiatives. To address this gap, the authors introduce the design process of OpenHealth Lake, a data management prototype platform based on a data lakehouse architecture, data federation, and the FAIR principles. The platform is designed using open-source tools, guided by system requirements identified in previously published studies and complemented by insights from the existing literature. The goal is to provide a flexible and complementary approach that allows organisations to customise data management systems to their specific requirements and resources, including cloud-based or self-hosted storage choices.

Innovation

The user study demonstrated that the OpenHealth Lake prototype is both usable and useful. Participants with varying technical backgrounds were able to interact with the platform through multiple interfaces: a user-friendly website, an open API, and Python and R packages. The results indicate that the platform successfully addresses the need for a unified data management solution in collaborative global health initiatives. The prototype's design, based on a data lakehouse architecture, data federation, and FAIR principles, proved to be adaptable and scalable. The study confirmed that the platform can be used by any organisation, regardless of their technical resources, as it supports both cloud-based and self-hosted storage choices. The findings highlight the potential of the lakehouse system to facilitate secure data storage and sharing, thereby supporting impactful research. The evaluation also underscored the importance of open-source tools in creating reproducible and customisable data management systems. Overall, the results validate the design approach and suggest that OpenHealth Lake can serve as a flexible and complementary solution for health data management.
Bioinformatics and health sciences continuously generate extensive heterogeneous datasets, making data management a complex challenge. In collaborative global health initiatives, secure storage and sharing of data are crucial to support impactful research. However, the absence of a unified data management platform complicates efficient data exchange and governance within these initiatives. To address this gap, the authors introduce the design process of OpenHealth Lake, a data management prototype platform based on a data lakehouse architecture, data federation, and the FAIR principles. The platform is designed using open-source tools, guided by system requirements identified in previously published studies and complemented by insights from the existing literature. The goal is to provide a flexible and complementary approach that allows organisations to customise data management systems to their specific requirements and resources, including cloud-based or self-hosted storage choices.
The design process of OpenHealth Lake followed a systematic approach. First, system requirements were identified from previously published studies and complemented by insights from the existing literature. These requirements guided the selection of open-source tools and the overall architecture. The platform was built as a prototype comprising a user-friendly website, an open API, and Python and R packages, allowing users to interact with the platform in multiple ways. The architecture is based on a data lakehouse model, which combines the scalability of data lakes with the structured querying capabilities of data warehouses. Data federation is employed to enable seamless access to distributed data sources without centralising storage, while the FAIR principles (Findable, Accessible, Interoperable, Reusable) ensure that data management practices support open science and reproducibility. To evaluate the prototype, a user study was conducted with participants having varying technical backgrounds. The study assessed both usability and usefulness of the platform. The evaluation aimed to determine whether the proposed data management prototype meets the needs of diverse users in health research settings. The prototype design showcases adaptability, scalability, and reproducibility, making it suitable for any organisation. The methodology emphasises flexibility, allowing organisations to customise the system to their specific requirements and resources, including cloud-based or self-hosted storage choices.

Why it matters

The OpenHealth Lake prototype addresses critical challenges in data management for bioinformatics and global health by integrating a data lakehouse architecture with data federation and FAIR principles. This combination allows for efficient data exchange and governance while maintaining flexibility for organisations with different resources. The use of open-source tools ensures reproducibility and adaptability, enabling organisations to customise the system to their specific needs. The user study, which included participants with varying technical backgrounds, confirmed the platform's usability and usefulness, demonstrating its potential for broad adoption. The design's support for both cloud-based and self-hosted storage options makes it accessible to a wide range of organisations, from those with extensive cloud infrastructure to those preferring on-premises solutions. The prototype's scalability and adaptability suggest it can handle the growing volume and heterogeneity of health data. However, the study also implies that further development and testing may be needed to fully realise the platform's potential in diverse real-world settings. The authors position OpenHealth Lake as a flexible and complementary approach that can be integrated with existing systems, rather than a one-size-fits-all solution. This aligns with the FAIR principles, promoting findable, accessible, interoperable, and reusable data. The discussion highlights the importance of user-friendly interfaces and multiple access methods (website, API, Python/R packages) to accommodate different user preferences and technical skills. The platform's architecture, illustrated in the Mermaid diagram below, shows the flow from data sources through federation and lakehouse storage to user interfaces. The design process, guided by system requirements and literature, ensures that the platform meets the needs of collaborative global health initiatives. Future work may involve expanding the user study and incorporating additional features based on feedback. Overall, OpenHealth Lake represents a significant step towards unified data management in health research, with the potential to enhance data sharing and collaboration across organisations.

Who should read this

CS practitioners and researchers

Opening member contentโ€ฆ