arXiv:cs.CL· Xilin Jiang, Shun Zhang, Tejas Jayashankar, Yinghao Aaron Li, Osama Hanna·· 3 小时前
对话语音美学模型 CVAM:用人类听者反馈强化学习描述语音美学
Conversational Voice Aesthetic Model with Reinforcement Learning from Human Listeners
AI 导读
研究提出对话语音美学模型 CVAM,一个用于在自然对话语境中描述真实或合成语音响应美学特征的语音大语言模型。团队基于 CANDOR 语料库构建 3k 条真实与合成响应、每条约 10 份人工标注的数据,经监督微调后用 Group Relative Policy Optimization 基于人类判断优化。
正文
Abstract:We introduce Conversational Voice Aesthetic Model, a speech large language model for describing the voice aesthetics of real or synthetic speech responses in natural conversational contexts. Given a context and a response speech, CVAM describes salient moments that characterize the voice and predicts nine categorical attributes spanning gender, pitch, pacing, emotion, and delivery. The key challenge lies in perceptual fields such as emotion and delivery, which are inherently subjective and lack definitive ground truth. Therefore, we collect ~10 human annotations for each of 3k real and synthetic responses derived from the CANDOR corpus. CVAM is supervised finetuned on synthesized aesthetic descriptions and labels, then optimized with Group Relative Policy Optimization on human judgments. Experiments show that CVAM better agrees with human listeners than Gemini 3.1 Pro and open-source speech LLMs, and outperforms single-human-vs.-rest agreement. Together, we demonstrate the importance of grounding voice aesthetics in human perception and propose a principled framework for human alignment.
| Subjects: | Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Machine Learning (cs.LG); Multimedia (cs.MM); Sound (cs.SD) |
| Cite as: | arXiv:2610.10868 [eess.AS] |
| (or arXiv:2610.10868v1 [eess.AS] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10868 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Xilin Jiang [view email]
[v1]
Wed, 7 Oct 2026 20:13:03 UTC (860 KB)
来源:arXiv:cs.CL · arxiv.org