arXiv:cs.LG· Afsara Benazir, Darius P\'etermann, Felix Xiaozhu Lin, Salar Rahili·· 4 小时前AI 评分38
SteerSpeech:用激活引导控制生成语音中的情感
Steerspeech: Activation Steering For Emotion Control In Generated Speech
AI 导读
SteerSpeech 是一个轻量级激活引导框架,通过向 TTS 模型的隐藏激活注入引导向量实现推理时情感控制,训练时冻结 TTS 主干,仅用低秩变换和多专家目标学习各情感的引导方向。在 Qwen3-TTS 上,该方法在已见、未见及带口音说话人上均实现更连续的情感控制,目标情感得分达基线的 1.08x-7.12x,高引导强度下说话人身份保持为 1.43x-1.46x。
正文
Abstract:Pretrained text-to-speech (TTS) models can generate expressive speech, but reliable inference-time emotion control remains challenging: prompts and reference audio offer coarse, inconsistent control, whereas specialized conditioning and model adaptation require costly training. We present SteerSpeech, a lightweight activation-steering framework that controls emotion by injecting steering vectors into hidden activations. For each target emotion we train a lightweight low-rank transform, using a multi-expert objective that encourages monotonic emotion control while preserving speaker identity and linguistic content, constraining steering drift, and keeping the TTS backbone frozen. To optimize through discrete speech tokens, we introduce a two-pass generation-and-replay pipeline using a straight-through estimator to backpropagate expert supervision through sampled tokens. At inference, a target-emotion steering direction is optimized with its respective transform and injected into the base TTS model. Objective and subjective evaluations with Qwen3-TTS across seen, unseen, and accented speakers show stronger continuous emotion control with limited speaker and content degradation. SteerSpeech achieves 1.08x-7.12x baseline target-emotion scores and for a representative emotion subjectively, it receives 78.1%-96.8% intensity preference and 1.43x-1.46x speaker-identity preservation at high steering strengths.
| Comments: | Under review at IEEE ICASSP 2027. This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible |
| Subjects: | Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS) |
| Cite as: | arXiv:2610.10415 [cs.SD] |
| (or arXiv:2610.10415v1 [cs.SD] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10415 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Afsara Benazir [view email]
[v1]
Wed, 7 Oct 2026 16:56:56 UTC (788 KB)
来源:arXiv:cs.LG · arxiv.org