arXiv:cs.AI· Mohammad Nur Hossain Khan, Subrata Biswas, Bashima Islam·· 4 小时前AI 评分36
DriftTTS:无需蒸馏的少步文本转语音,基于分布匹配漂移
DriftTTS: Few-Step Text-to-Speech Without Distillation via Distribution-Matching Drift
AI 导读
DriftTTS 是一种无需生成式教师模型、蒸馏或对抗判别训练的少步梅尔频谱生成器,采用分布匹配漂移目标,在原始梅尔与冻结掩码自编码器编码器定义的梅尔域特征空间中训练。
正文
Abstract:Few-step neural text-to-speech models often rely on short- ened diffusion or flow-matching schedules, or on distillation from pretrained multi-step teachers. To avoid these depen- dencies, we present DriftTTS, a few-step mel-spectrogram generator trained without a generative teacher, distillation, or adversarial discrimination. DriftTTS uses a distribution- matching drift objective in a mel-domain feature space defined by raw mels and a frozen masked-autoencoder encoder pretrained on the same LJSpeech training split. On-policy rollout trains the decoder on its own interme- diate states and supports inference up to the trained roll- out depth. On LJSpeech, DriftTTS at NFE=4 achieves 3.87 dB MCD and 3.7% WER, compared with 3.85 dB and 3.4% for Matcha-TTS. In a fully paired blind listen- ing test, DriftTTS obtains 4.18 MOS, compared with 3.96 for Matcha-TTS and 4.22 for ground truth. These results demonstrate competitive few-step synthesis without a pre- trained generative teacher. Code can be found at https: //github.com/BASHLab/driftTTS.git
| Subjects: | Sound (cs.SD); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.03390 [cs.SD] |
| (or arXiv:2610.03390v1 [cs.SD] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03390 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Mohammad Nur Hossain Khan [view email]
[v1]
Fri, 2 Oct 2026 14:38:59 UTC (304 KB)
来源:arXiv:cs.AI · arxiv.org