跳到正文
arXiv:cs.LG· Taejun Kim, Wonil Kim, Jongmin Jung, Hyeongseok Wi, Sangeun Kum, Keunhyoung Luke Kim, Taehyoung Kim, Dongjoo Moon, Seungsoon Park, Taewan Kim, Virginie Berger, Juhan Nam, Jongpil Lee·· 4 小时前AI 评分37

如何验证 AI 生成音乐的归属:用输入溯源与输出审核追踪 MixAudio 与 musicDNA

Tracing Inputs, Verifying Outputs: Validating Attribution in Music Generation

AI 导读

该论文提出以输入溯源验证 AI 音乐生成中音频来源及其对输出的影响,生成器 MixAudio 仅以音频为条件、不含文本输入。提示词遵循测试与受控输入替换显示,生成的 stems 在音色上跟随提示音频、在和声上跟随上下文音频;随后用音乐版本识别模型 musicDNA 审核记忆化,发现输入记录之外的重现很少,在人工判定样本上精确率与召回率均高于其他被测检测器。

正文

Authors:Taejun Kim, Wonil Kim, Jongmin Jung, Hyeongseok Wi, Sangeun Kum, Keunhyoung Luke Kim, Taehyoung Kim, Dongjoo Moon, Seungsoon Park, Taewan Kim, Virginie Berger, Juhan Nam, Jongpil Lee

View PDF HTML (experimental)

Abstract:How can we verify whose music contributed to an AI-generated output? This paper demonstrates how input-based attribution can provide verifiable evidence of which audio sources were used in a generation and whether they shaped the output. To do so, we condition the generation solely on audio without any text input, then trace the inputs behind each output, and establish their musical effect. In prompt adherence tests and controlled input swaps, the stems generated by our generator, MixAudio, follow the prompt audio in timbre and the context audio in harmony. Yet these outputs may still reproduce training data not supplied as inputs. We therefore audit memorization with our musical version identification model, musicDNA, and find few reproductions outside the input records. On human-judged cases within the flagged pool, it achieves higher precision and recall than the other tested memorization detectors. The two evaluations suggest that input records and output analysis provide complementary evidence for attribution, on which rights-holder reporting and compensation can draw as the AI music economy takes shape. Audio examples are available at this https URL
Comments: 15 pages, 4 figures. Audio examples: this https URL
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
Cite as: arXiv:2610.09637 [cs.SD]
  (or arXiv:2610.09637v1 [cs.SD] for this version)
  https://doi.org/10.48550/arXiv.2610.09637

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Taejun Kim [view email]
[v1] Wed, 7 Oct 2026 08:14:19 UTC (629 KB)

来源:arXiv:cs.LG · arxiv.org