跳到正文
arXiv:cs.LG· Yudi Zhang, Mingyu Cao, Lu Yin, Mykola Pechenizkiy, Shiwei Liu·· 4 小时前

DataSense-Bench:迈向 AI 科学家的重要一步

DataSense-Bench: The First Step Toward an AI Scientist

AI 导读

研究者提出 DataSense-Bench,用数据选择与性能预测任务检验前沿 AI 模型是否具备"数据感"。智能体可查看数据、编写执行分析代码、运行模型前向,但不能训练模型或接触真实评测任务;基准基于 OpenThoughts-Agent 与 EnvScaler 的轨迹,分别在 TBLite 和 BFCL 上评测。

正文

View PDF HTML (experimental)

Abstract:As claims about recursive self-improvement (RSI) and artificial general intelligence (AGI) proliferate, we ask a simple question: do frontier AI models have a sense of data, i.e., can they reliably select the right data for training? We introduce DataSense-Bench to study this capability through the fundamental problem of data selection and performance forecasting in machine learning. We ask AI agents to select and rank candidate training subsets that can be used to fine-tune a small LLM model. Agents are allowed to inspect the data, write and execute analysis code, and run model forward passes, but can not train the model or access the actual evaluation tasks. We then fine-tune the base model on each selected subset and evaluate its post-training performance under a standardized protocol. We instantiate the benchmark in terminal problem solving and tool use, selecting trajectories from OpenThoughts-Agent and EnvScaler and evaluating on TBLite and BFCL, respectively. We then evaluate the agents along two complementary dimensions: the post-training performance of the top-ranked subset, reflecting the ability to identify high-value training data, and ranking accuracy, reflecting the ability to predict the relative performance of the selected subsets. In our experiments, selection gains over random selection are limited; agents do not reliably rank their selected groups, and ranking ability does not hold consistently across tasks: Astra identifies the best group in all three tool-use runs but in only one of three terminal runs. Analysis of execution traces on both tasks shows that agents often use similar data signals while interpreting their training value differently.
Comments: 25 pages. Project page: this https URL . Code: this https URL
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.12190 [cs.LG]
  (or arXiv:2610.12190v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.12190

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yudi Zhang [view email]
[v1] Thu, 8 Oct 2026 15:50:57 UTC (191 KB)

来源:arXiv:cs.LG · arxiv.org