arXiv:cs.LG(机器学习,全量分类)· Sidi Chang, Peiying Zhu·· 9 小时前AI 评分28
合成语音研究对象的生成溯源审计:NeurIPS 数据归因工作坊论文
Generation Provenance Before Behavior Attribution: Auditing Synthetic Speech Research Objects
AI 导读
研究者提出"生成溯源"基底,将合成研究对象的源规格、生成内容、波形、目标、事实要求、质量信号、审核谱系与不可变清单身份绑定,并在一个私有日语护理交接流水线中审计。
正文
Abstract:Attributing model behavior to synthetic training data requires knowing what produced each training item before estimating what that item caused. A waveform-label pair does not preserve this knowledge. We propose a generation-provenance substrate in which a synthetic research object binds source specification, generated content, waveform, target, fact requirements, quality signals, review lineage, and immutable manifest identity. Producer and selection mechanism determine evidentiary meaning; storage location and variable name do not. We audit this substrate in a private Japanese care-handoff pipeline. A 113-asset review population contains 1.552 hours of synthetic speech across six scenario families; all items have linked audio, transcripts, candidate notes, and fact checklists, but human evidence is selective and source-specific. Two faithful-only manifests are scenario-seed-disjoint and immutably versioned, while exact upstream attribution remains blocked by floating generator aliases, missing per-clip TTS and code stamps, and an unversioned checking prompt. We argue that generation provenance is necessary but not sufficient for behavior attribution: it defines the candidate causal graph and audit units, whereas contributive attribution still requires frozen training runs and intervention or influence evidence. The paper contributes a compact provenance contract, an audit protocol, and a bounded case study for synthetic-data attribution; controlled research access may be offered, but we do not claim causal training-data attribution, clinical validity, or unrestricted public release.
| Comments: | Accepted to the Third NeurIPS Workshop on Attributing Model Behavior at Scale: Data Attribution and Provenance. 4 pages, 0 figures, 1 table. An aggregate reproducibility package is available from the authors on request! |
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG) |
| MSC classes: | 68T05 |
| ACM classes: | I.2.6 |
| Cite as: | arXiv:2610.01378 [cs.AI] |
| (or arXiv:2610.01378v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.01378 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Sidi Chang [view email]
[v1]
Thu, 1 Oct 2026 09:45:00 UTC (10 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org