跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Joss Armstrong·· 1 天前AI 评分42

来源识别不等于适应度测试:合成数据归因的极限

Source Identification Is Not Fitness Testing: Measuring the Limits of Synthetic-Data Attribution

AI 导读

一项基于金融风险文本的研究发现,生成段落的生成器归因在原始文本上准确率为 98.7%,但改写后降至 53.1%,风格重写后仅 29.0%。在生成与重训练的三轮对比中,使用来源信息与使用参考模型评分的两种筛选规则选出了不同样本,但未检测到所得模型退化存在稳定差异。结果表明,识别数据来源与识别哪些数据对训练有用是两个独立问题。

正文

View PDF HTML (experimental)

Abstract:Repeated training on model-generated data can degrade later models. One possible response is to use provenance when deciding which generated examples to reuse. We test both how reliably that provenance can be recovered and whether it helps identify better training data. Using financial-risk text, we first identify the source of generated passages and then repeat the test after rewriting them. Generator attribution is 98.7% accurate on the original passages but falls to 53.1% after paraphrasing and 29.0% after style rewriting. Generated-versus-human detection remains close to perfect against the tested human comparison set. We then compare two ways of selecting generated examples over three rounds of generation and retraining. One uses source information. The other uses a score from a separate reference model. The two rules select different examples, but the planned comparison does not detect a stable difference in the degradation of the resulting models. The results show that identifying where data came from and identifying which data are useful for training are separate problems. The experiment therefore separates source identity, criterion-facing selection, and recursive training outcome: neither the provenance score nor the tested criterion-facing proxy is established as sufficient for future recursive behaviour.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.00417 [cs.LG]
  (or arXiv:2610.00417v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.00417

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Joss Armstrong [view email]
[v1] Wed, 30 Sep 2026 14:34:55 UTC (46 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org