跳到正文
arXiv:cs.LG· Berke Arda, Ahmetcan Yavuz, Paul Gerry, Sebastian Lobentanzer, Nobin Sarwar, Joan Giner-Miguelez, Kongtao Chen, Luyao Zhang, Mrinmaya Sachan, Mubashara Akhtar·· 4 小时前AI 评分42

CroissantMiner:为 ML 数据集自动提取与验证 Croissant 元数据

CroissantMiner: Automated Extraction and Validation of Croissant Metadata for ML Datasets

AI 导读

研究者发布首个端到端评测 Croissant 元数据提取的基准 CroissantMiner,包含 602 篇论文(102 篇人工验证金标注、500 篇 LLM 银标注),覆盖核心与 Responsible AI(RAI)字段,并被 NeurIPS 2026 接收。

正文

View PDF HTML (experimental)

Abstract:Croissant has emerged as a standard for machine-readable dataset metadata, yet populating its fields remains labor-intensive and requires careful reading of accompanying dataset documentation. We present the first benchmark enabling end-to-end evaluation of metadata extraction aligned with a community-standard schema. The benchmark comprises 602 papers, including 102 with human-validated gold annotations and 500 with LLM-generated silver annotations, covering the full Croissant schema with both core and Responsible AI (RAI) fields. Using this benchmark, we evaluate a range of extraction systems spanning frontier models, open-weight models, and agentic architectures, under a two-tier evaluation framework that combines rule-based scoring with an LLM judge selected via human audit. We find that single-pass extraction consistently outperforms the four agentic architectures we evaluate: across backbones, these decomposed variants achieve lower accuracy than a single full-context pass. The largest gap appears on long-form RAI fields, which require synthesizing and interpreting information scattered across a paper rather than copying it from a single location, a setting where current systems remain far from reliable. We release the benchmark, evaluation code, judge audit, a live demo, and a leaderboard open to new systems.
Comments: Accepted at NeurIPS 2026 (Track on Evaluations and Datasets). Website: this https URL
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Machine Learning (cs.LG)
Cite as: arXiv:2610.07132 [cs.CL]
  (or arXiv:2610.07132v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.07132

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Berke Arda [view email]
[v1] Mon, 5 Oct 2026 17:59:48 UTC (1,403 KB)

来源:arXiv:cs.LG · arxiv.org