跳到正文
arXiv:cs.AI· Nityanand Mathur, Hamees Sayed, Ayush Pratap Singh·· 5 小时前AI 评分38

精修提升可懂度,搜索提升身份:掩码扩散 TTS 中测试时计算买到了什么

Refinement Buys Intelligibility, Search Buys Identity: What Test-Time Compute Buys in Masked-Diffusion TTS

AI 导读

研究训练了 15 个掩码扩散 codec TTS 模型(19-133M 参数、3 个随机种子),在 2000 小时语音上考察推理精修步数 T∈[1,16] 的作用:精修弥补了 86.2% 的可懂度差距,但只弥补 46.4% 的说话人身份差距,呈 1.86 倍不对称。

正文

View PDF HTML (experimental)

Abstract:Diffusion language models for text-to-speech combine two forms of computation: model depth (parameters) and refinement steps (inference budget). We ask whether they scale equally across capabilities. We train 15 masked-diffusion codec TTS models varying depth (19-133M parameters, 3 seeds) on 2,000 hours of speech and sweep refinement steps T in [1,16] at inference, measuring zero-shot synthesis via ASR word error rate (intelligibility) and speaker verification (identity) on 174 held-out speakers. Against measured floors, refinement closes 86.2% of the intelligibility range but only 46.4% of the identity range - a 1.86x asymmetry robust across multiple error metrics. Retraining at 3x and 6x schedule attenuates but does not reverse this gap (1.84 to 1.36 to 1.23x), because intelligibility saturates with steps while identity continues improving. Best-of-K search recovers speaker identity where refinement fails, with 64.6-79.0% win rates across four independent encoders. Depth and steps are not interchangeable: separable B(d)B(T) fits significantly better (Delta AICc=+69.3) than substitution models. Analysis shows 62% of remaining identity deficit lies in the codec, not the generator. We conclude that refinement and depth target different bottlenecks and should be optimized separately.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.03320 [cs.AI]
  (or arXiv:2610.03320v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.03320

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Nityanand Mathur Mr [view email]
[v1] Fri, 2 Oct 2026 13:56:08 UTC (152 KB)

来源:arXiv:cs.AI · arxiv.org