跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Sidi Chang, Peiying Zhu·· 9 小时前AI 评分28

合成语音研究对象的生成溯源审计:NeurIPS 数据归因工作坊论文

Generation Provenance Before Behavior Attribution: Auditing Synthetic Speech Research Objects

AI 导读

研究者提出"生成溯源"基底,将合成研究对象的源规格、生成内容、波形、目标、事实要求、质量信号、审核谱系与不可变清单身份绑定,并在一个私有日语护理交接流水线中审计。

正文

View PDF HTML (experimental)

Abstract:Attributing model behavior to synthetic training data requires knowing what produced each training item before estimating what that item caused. A waveform-label pair does not preserve this knowledge. We propose a generation-provenance substrate in which a synthetic research object binds source specification, generated content, waveform, target, fact requirements, quality signals, review lineage, and immutable manifest identity. Producer and selection mechanism determine evidentiary meaning; storage location and variable name do not. We audit this substrate in a private Japanese care-handoff pipeline. A 113-asset review population contains 1.552 hours of synthetic speech across six scenario families; all items have linked audio, transcripts, candidate notes, and fact checklists, but human evidence is selective and source-specific. Two faithful-only manifests are scenario-seed-disjoint and immutably versioned, while exact upstream attribution remains blocked by floating generator aliases, missing per-clip TTS and code stamps, and an unversioned checking prompt. We argue that generation provenance is necessary but not sufficient for behavior attribution: it defines the candidate causal graph and audit units, whereas contributive attribution still requires frozen training runs and intervention or influence evidence. The paper contributes a compact provenance contract, an audit protocol, and a bounded case study for synthetic-data attribution; controlled research access may be offered, but we do not claim causal training-data attribution, clinical validity, or unrestricted public release.
Comments: Accepted to the Third NeurIPS Workshop on Attributing Model Behavior at Scale: Data Attribution and Provenance. 4 pages, 0 figures, 1 table. An aggregate reproducibility package is available from the authors on request!
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
MSC classes: 68T05
ACM classes: I.2.6
Cite as: arXiv:2610.01378 [cs.AI]
  (or arXiv:2610.01378v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.01378

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Sidi Chang [view email]
[v1] Thu, 1 Oct 2026 09:45:00 UTC (10 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org