arXiv:cs.CL· Zirui Li, Rech Silas, Lauri Juvela, Tom Backstrom, Mikko Kurimo·· 3 小时前AI 评分32
EmphTTS:用强化学习实现词级重音控制的 TTS 系统
EmphTTS: an emphasis-control TTS with reinforcement learning
AI 导读
EmphTTS 是一个非自回归 TTS 系统,将 GRPO 应用于时长预测器并配合重音定位奖励,直接优化词级重音控制。评测显示其在重音可控性和客观指标上均表现最佳,主观偏好测试中显著优于合成 groundtruth 及多数基线。消融实验表明 GRPO 的重音实现效果超过监督微调时长建模和简单语速调整,同时缓解了独立训练的时长预测器与 TTS 模型之间的不匹配。
正文
Abstract:Generating controllable and human-like emphasis remains an open challenge in text-to-speech, even when explicit emphasis control signals are provided in the text input, limiting the communicative accuracy of synthetic speech in real-world applications. Reinforcement learning has recently shown promise for post-training TTS systems to align with human preference, yet existing methods have not been applied to word-level prosodic control. We present EmphTTS, a non-autoregressive TTS system that applies Group Relative Policy Optimization (GRPO) to the duration predictor with an emphasis localization reward, enabling direct optimization for word-level emphasis. Evaluations show that EmphTTS achieves the best emphasis controllability and performs the best in emphasis objective evaluation. In subjective preference tests, EmphTTS is significantly preferred over synthetic groundtruth and most baselines. Ablation studies show that GRPO improves emphasis realization beyond supervised-finetuning-based duration modeling and simple speaking-rate adjustment, while alleviating the mismatch between the independently trained duration predictor and TTS model.
| Comments: | 5 pages. Submitted to ICASSP 2027 |
| Subjects: | Audio and Speech Processing (eess.AS); Computation and Language (cs.CL) |
| Cite as: | arXiv:2609.27599 [eess.AS] |
| (or arXiv:2609.27599v2 [eess.AS] for this version) | |
| https://doi.org/10.48550/arXiv.2609.27599 arXiv-issued DOI via DataCite |
Submission history
From: Zirui Li [view email]
[v1]
Wed, 23 Sep 2026 09:15:25 UTC (1,776 KB)
[v2]
Tue, 6 Oct 2026 08:34:59 UTC (1,776 KB)
来源:arXiv:cs.CL · arxiv.org