Jadwal Sholat

Memuat jadwal sholatโ€ฆ

Computer Science editorial

Open AccessOA2026

Croissant Baker: Metadata Generation for Discoverable, Governable, and Reusable ML Datasets

A local-first, open-source CLI tool that generates validated Croissant metadata directly from dataset directories, achieving 97โ€“100% agreement with ground truth across 140+ datasets.
Rafi Al Attrach; Rajna Fani; Sebastian Lobentanzer; Joan Giner-Miguelez; Debanshu Das; Varuni H. K.; Nobin Sarwar; Rajat Ghosh; Anwai Archit; Surbhi Motghare; Christina Conrad Parry; Luis Oala; Lara Grosso; Joaquin Vanschoren; Steffen Vogler; Sujata Goswami; Eric S. Rosenthal; Marzyeh Ghassemi; Matthew McDermott; Tom Pollardยท 2026ยท DOI 10.48550/arXiv.2605.15079

The core problem

Machine learning datasets increasingly require standardized metadata to support discovery, automated ingestion, and reproducible analysis. Croissant has emerged as the metadata standard for ML datasets, providing a structured, JSON-LD-based format that makes dataset properties machine-checkable across platforms. Adoption has accelerated, and NeurIPS now requires Croissant metadata in every submission to its dataset tracks. However, in practice, Croissant generation usually starts with uploading data to a public platformโ€”a path infeasible for governed and large local repositories that hold much of the high-value data ML increasingly relies on. This paper introduces Croissant Baker, a local-first, open-source command-line tool that generates validated Croissant metadata directly from a dataset directory through a modular handler registry. The authors evaluate Croissant Baker on over 140 datasets, scaling to MIMIC-IV at 886 million rows and 374 Parquet files. On held-out comparisons against producer-authored or standards-derived ground truth, Croissant Baker reaches 97โ€“100% agreement across multiple domains.

Innovation

Croissant Baker was evaluated on over 140 datasets, demonstrating its ability to handle diverse data types and scales. The largest dataset tested was MIMIC-IV, a critical care database containing 886 million rows and 374 Parquet files. Despite the scale, Croissant Baker successfully generated validated Croissant metadata.

On held-out comparisons against producer-authored or standards-derived ground truth, Croissant Baker achieved 97โ€“100% agreement across multiple domains. This high level of agreement indicates that the automatically generated metadata is nearly indistinguishable from manually curated or standards-compliant metadata. The evaluation covered domains such as healthcare, natural language processing, and computer vision, showcasing the tool's versatility.

The results are summarized in the following table:

| Dataset Scale | Number of Datasets | Agreement Range |
|---------------|-------------------|-----------------|
| Small to Medium | 140+ | 97โ€“100% |
| Large (MIMIC-IV) | 1 | 97โ€“100% |

These findings suggest that Croissant Baker can reliably automate metadata generation for ML datasets, reducing the burden on dataset producers and enabling compliance with emerging s

Machine learning datasets increasingly require standardized metadata to support discovery, automated ingestion, and reproducible analysis. Croissant has emerged as the metadata standard for ML datasets, providing a structured, JSON-LD-based format that makes dataset properties machine-checkable across platforms. Adoption has accelerated, and NeurIPS now requires Croissant metadata in every submission to its dataset tracks. However, in practice, Croissant generation usually starts with uploading data to a public platformโ€”a path infeasible for governed and large local repositories that hold much of the high-value data ML increasingly relies on. This paper introduces Croissant Baker, a local-first, open-source command-line tool that generates validated Croissant metadata directly from a dataset directory through a modular handler registry. The authors evaluate Croissant Baker on over 140 datasets, scaling to MIMIC-IV at 886 million rows and 374 Parquet files. On held-out comparisons against producer-authored or standards-derived ground truth, Croissant Baker reaches 97โ€“100% agreement across multiple domains.
Croissant Baker is designed as a local-first tool that operates directly on a dataset directory without requiring data upload to a public platform. Its architecture is modular, centered on a handler registry that dispatches metadata extraction tasks to specialized handlers based on file format and dataset characteristics. This design enables extensibility and adaptation to diverse data types and structures.

Why it matters

The introduction of Croissant Baker addresses a critical gap in the ML data ecosystem: the need for metadata generation that works in governed and local environments. By operating locally, it avoids the privacy, security, and logistical issues associated with uploading sensitive data to public platforms. This is particularly important for high-value datasets in healthcare, finance, and other regulated domains.

The modular handler registry design provides flexibility and extensibility. As new data formats emerge, the community can contribute handlers, ensuring that Croissant Baker remains relevant. The high agreement rates (97โ€“100%) with ground truth demonstrate that automated metadata generation can meet the quality standards required for machine-checkable metadata.

However, the evaluation is limited to the datasets tested, and further validation on a broader range of domains and formats would strengthen the findings. Additionally, the tool's performance on extremely large datasets beyond MIMIC-IV remains to be explored. Future work could focus on optimizing performance for distributed storage systems and integrating with data versioning tools.

Overall, Croissant Baker represents a significant step toward making ML datasets more discoverable, governable, and reusable by automating the generation of standardized metadata in a local-first manner.

Who should read this

CS practitioners and researchers

Opening member contentโ€ฆ