跳到正文
arXiv:cs.CL· Shiao Zhu, Lianbo Liu, Sizhen Lyu, Yuzhe Wang, Sheng Li, Takahiro Shinozaki·· 3 小时前

SRSP:用语音奖励驱动的风格规划改进对话式 TTS

Beyond Speech Captions: Speech-Rewarded Style Planning for Conversational Text-to-Speech

AI 导读

研究提出 Speech-Rewarded Style Planning(SRSP),通过冻结的下游 TTS 模型训练基于文本的风格规划器,以目标语音 token 的 teacher-forced 似然为奖励,采用 GRPO 优化。

正文

View PDF HTML (experimental)

Abstract:Natural-language style descriptions provide an interpretable interface between large language models (LLMs) and controllable text-to-speech (TTS). However, using descriptions as pseudo-labels compresses target acoustics into text, and descriptive fidelity need not imply effective control of a particular synthesizer. We empirically show that speech-text alignment only weakly predicts downstream acoustic similarity among candidate instructions for the same utterance. We therefore propose Speech-Rewarded Style Planning (SRSP), which trains a text-based style planner through a frozen downstream TTS model. Given dialogue history and response text, the planner generates candidate instructions and is optimized with group-relative policy optimization (GRPO), using the teacher-forced likelihood of target speech tokens as the reward. On an English subset of the ISCSLP 2026 CoT-TTS corpus, SRSP achieves higher speech-style and emotion similarity to target speech and lower mel-cepstral distortion than the Base LLM and target-audio-informed captioning baselines. LLM-based expressive speech evaluation further shows gains over all baselines in contextual appropriateness and reference consistency.
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
Cite as: arXiv:2610.11461 [cs.SD]
  (or arXiv:2610.11461v1 [cs.SD] for this version)
  https://doi.org/10.48550/arXiv.2610.11461

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Shiao Zhu [view email]
[v1] Thu, 8 Oct 2026 08:12:28 UTC (92 KB)

来源:arXiv:cs.CL · arxiv.org