AutoDataBench:加速自动研究的数据中心化测试平台
AutoDataBench: A Data-centric Testbed for Accelerating Auto Research
研究者提出 AutoDataBench,一个隔离数据因素、固定训练框架与算力等非数据变量的受控测试平台,用于评估 LLM 的数据智能。该平台围绕数据诊断、数据组织与数据构建三类任务,在工具调用、检索和知识注入场景下考察前沿 LLM 通过迭代实验改进训练数据的能力,并对比训练前预测与实际结果以检验其数据效应推理。复用 AutoDataBench 轨迹进行中期训练可提升下游编码性能。
Published on Sep 30
Authors:
,
,
,
,
,
,
,
,
,
,
Abstract
Existing auto-research benchmarks often entangle multiple sources of improvement, including training frameworks, hyperparameters, compute budgets, and data, making it difficult to attribute why one frontier agent outperforms another to specific research capabilities. In this work, we isolate and systematically evaluate Data Intelligence: an agent's ability to understand, manipulate, and improve the data that shapes model capabilities. We introduce AutoDataBench, a controlled testbed built on a conceptual framework of data intelligence spanning data diagnosis, data organization, and data construction, instantiated through three highly curated optimization tasks while holding non-data factors fixed. Across tool use, retrieval, and knowledge injection, we evaluate frontier LLMs' ability to improve training data through iterative experimentation under task-specific resource budgets. Beyond optimization performance, we ask: do LLMs understand what their data interventions do? We compare predictions made before training with observed outcomes to seek evidence of data-effect reasoning beyond trial and error, and explore whether iterative feedback helps LLMs better understand how changes to training data affect model performance. Finally, we show that reusing AutoDataBench trajectories for mid-training improves downstream coding performance, highlighting its value in both evaluating data intelligence and generating high-quality training data. Code and resources are available at https://github.com/AutoDataBench/AutoDataBench.
View arXiv page View PDF GitHub 3 Add to collection
Get this paper in your agent:
hf papers read 2609.40097
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash
Models citing this paper 3
Updated about 10 hours agoAutoDataBench/Retrieval-resources
Question Answering • Updated about 10 hours agoAutoDataBench/Knowledge-Injection-resources
Updated about 10 hours agoAutoDataBench/Function-Calling-resources
Datasets citing this paper 0
No dataset linking this paper
Cite arxiv.org/abs/2609.40097 in a dataset README.md to link it from this page.
Spaces citing this paper 0
No Space linking this paper
Cite arxiv.org/abs/2609.40097 in a Space README.md to link it from this page.
Collections including this paper 0
No Collection including this paper
Add this paper to a collection to link it from this page.
来源:HuggingFace Daily Papers(社区热门论文) · huggingface.co