跳到正文
arXiv:cs.CL· Pulak Kuli·· 3 小时前AI 评分39

Moshi 全双工语音模型的情感引导:情感方向由几何结构而非标签决定

Steering Follows Geometry, Not Labels: Emotion Directions in a Full-Duplex Speech Model

AI 导读

研究者在全双工语音语言模型 Moshi 上测试四种情感的平均差激活引导,该方法每帧仅需几次向量加法且无需重训练,情感可从其残差流中线性解码,但激活引导只能部分实现且效果不均。happy、angry、surprise 共享同一引导方向,sad 则具有独立可引导方向;三种情感共享的分量无法被同等从所有情感中投影消除。

正文

View PDF HTML (experimental)

Abstract:Full-duplex voice agents need to modulate emotion and delivery during real-time conversations, when de-escalating a complaint, carrying urgency in dispatch, softening a clinical result. Emotion and delivery control is well studied for TTS and turn based models through prompt-conditioned synthesis, reference-conditioned synthesis and activation steering; PersonaPlex controls identity in a duplex model but not affect.
We study emotion steering in Moshi, a fully open sourced full-duplex speech language model, across four emotions, using mean-difference activation steering, which costs only a few vector additions per frame and no retraining.
We show that emotion is linearly decodable from Moshi's residual stream, but activation steering is only partially achievable, and unevenly so; as happy, angry and surprise steer towards a shared direction while sad is distinctly steerable. We also show that the shared component across the three emotions cannot simply be projected away from all the emotions equally.
Comments: Accepted at the NeurIPS 2026 Workshop on Real-Time Conversational Agents (RTCA), Sydney. OpenReview: this https URL
Subjects: Computation and Language (cs.CL); Sound (cs.SD)
ACM classes: I.2.7; I.2.6
Cite as: arXiv:2610.08887 [cs.CL]
  (or arXiv:2610.08887v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.08887

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Pulak Kuli [view email]
[v1] Tue, 6 Oct 2026 15:29:37 UTC (53 KB)

来源:arXiv:cs.CL · arxiv.org