arXiv:cs.CL· Shiao Zhu, Lianbo Liu, Sizhen Lyu, Yuzhe Wang, Sheng Li, Takahiro Shinozaki·· 3 小时前
SRSP:用语音奖励驱动的风格规划改进对话式 TTS
Beyond Speech Captions: Speech-Rewarded Style Planning for Conversational Text-to-Speech
AI 导读
研究提出 Speech-Rewarded Style Planning(SRSP),通过冻结的下游 TTS 模型训练基于文本的风格规划器,以目标语音 token 的 teacher-forced 似然为奖励,采用 GRPO 优化。
正文
Abstract:Natural-language style descriptions provide an interpretable interface between large language models (LLMs) and controllable text-to-speech (TTS). However, using descriptions as pseudo-labels compresses target acoustics into text, and descriptive fidelity need not imply effective control of a particular synthesizer. We empirically show that speech-text alignment only weakly predicts downstream acoustic similarity among candidate instructions for the same utterance. We therefore propose Speech-Rewarded Style Planning (SRSP), which trains a text-based style planner through a frozen downstream TTS model. Given dialogue history and response text, the planner generates candidate instructions and is optimized with group-relative policy optimization (GRPO), using the teacher-forced likelihood of target speech tokens as the reward. On an English subset of the ISCSLP 2026 CoT-TTS corpus, SRSP achieves higher speech-style and emotion similarity to target speech and lower mel-cepstral distortion than the Base LLM and target-audio-informed captioning baselines. LLM-based expressive speech evaluation further shows gains over all baselines in contextual appropriateness and reference consistency.
| Subjects: | Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS) |
| Cite as: | arXiv:2610.11461 [cs.SD] |
| (or arXiv:2610.11461v1 [cs.SD] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11461 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Shiao Zhu [view email]
[v1]
Thu, 8 Oct 2026 08:12:28 UTC (92 KB)
来源:arXiv:cs.CL · arxiv.org