跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Shivalika Singh, Andrija Djurisic, Gbemileke Onilude, Sudip Roy, Sara Hooker·· 14 小时前AI 评分42

Invent-A-Dataset:零种子数据下测量数据集生成能力

Invent a Dataset: Measuring dataset generation abilities with zero seed

AI 导读

研究者提出 Invent-A-Dataset,一个基于提示词的系统,可从数据集描述直接生成真实、大规模的后训练数据集,面向完全没有目标能力数据的零数据场景。

正文

View PDF HTML (experimental)

Abstract:Building datasets remains one of the most manual and brittle parts of AI development. In this technical report, we focus on the most extreme but also most prevalent setting real world practitioners face: a zero data regime. Here, practitioners don't have any data for the capability they want to learn. We introduce Invent-A-Dataset which is a prompt based system to go from dataset description to realistic and large scale post-training datasets. We evaluate Invent-A-Dataset against five frontier model APIs including Anthropic, Google, Open AI, DeepSeek, Zai. Across eight task types and dataset sizes up to 20K samples, Invent-A-Dataset significantly outperforms with both the highest quality (17% relative gains) while simultaneously producing the most diverse samples (19% relative gains). Its diversity advantage widens with scale of training dataset size (from parity at 200 samples to 37% relative gains at 20K samples). This translates into considerable downstream training gains, resulting in far more performant post-trained models. Invent-A-Dataset fine-tune consistently ranks higher compared to other generator fine-tunes across different post-trained model architectures.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.01674 [cs.LG]
  (or arXiv:2610.01674v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.01674

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Andrija Djurisic [view email]
[v1] Thu, 1 Oct 2026 13:32:06 UTC (3,092 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org