arXiv:cs.CL· Pulak Kuli·· 3 小时前AI 评分39
Moshi 全双工语音模型的情感引导:情感方向由几何结构而非标签决定
Steering Follows Geometry, Not Labels: Emotion Directions in a Full-Duplex Speech Model
AI 导读
研究者在全双工语音语言模型 Moshi 上测试四种情感的平均差激活引导,该方法每帧仅需几次向量加法且无需重训练,情感可从其残差流中线性解码,但激活引导只能部分实现且效果不均。happy、angry、surprise 共享同一引导方向,sad 则具有独立可引导方向;三种情感共享的分量无法被同等从所有情感中投影消除。
正文
Abstract:Full-duplex voice agents need to modulate emotion and delivery during real-time conversations, when de-escalating a complaint, carrying urgency in dispatch, softening a clinical result. Emotion and delivery control is well studied for TTS and turn based models through prompt-conditioned synthesis, reference-conditioned synthesis and activation steering; PersonaPlex controls identity in a duplex model but not affect.
We study emotion steering in Moshi, a fully open sourced full-duplex speech language model, across four emotions, using mean-difference activation steering, which costs only a few vector additions per frame and no retraining.
We show that emotion is linearly decodable from Moshi's residual stream, but activation steering is only partially achievable, and unevenly so; as happy, angry and surprise steer towards a shared direction while sad is distinctly steerable. We also show that the shared component across the three emotions cannot simply be projected away from all the emotions equally.
| Comments: | Accepted at the NeurIPS 2026 Workshop on Real-Time Conversational Agents (RTCA), Sydney. OpenReview: this https URL |
| Subjects: | Computation and Language (cs.CL); Sound (cs.SD) |
| ACM classes: | I.2.7; I.2.6 |
| Cite as: | arXiv:2610.08887 [cs.CL] |
| (or arXiv:2610.08887v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08887 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Pulak Kuli [view email]
[v1]
Tue, 6 Oct 2026 15:29:37 UTC (53 KB)
来源:arXiv:cs.CL · arxiv.org