HuggingFace Daily Papers·· 1 天前AI 评分39
HiPLEX:面向全双工语音语言模型的分层策略因子化框架
HiPLEX: Hierarchical Policy Factorization for Full Duplex Speech Language Models
AI 导读
HiPLEX 是一个强化学习框架,将预训练的全双工文本策略因子化为决定"何时发声"的控制策略和决定"发什么"的条件内容策略,后者仅在选择 'con' 时激活 token。
正文
Abstract:As human--AI interactions become more conversational, full-duplex speech language models capable of natural real-time dialogue are growing in importance. Beyond generating appropriate responses, these models must coordinate turn-taking, backchanneling, and floor management in real time. Reinforcement learning (RL) provides a way to refine these behaviors through direct feedback on interaction outcomes. However, existing RL methods either apply timing feedback to a token policy or optimize semantic content, leaving the joint improvement of timing and content unresolved. We introduce HiPLEX, an RL framework that factorizes a pretrained full-duplex text policy into a control policy that decides when to emit content and a conditional content policy that decides what to emit. The first factor selects among 'pad', 'epad', and 'con'. The second selects a token only when 'con' is chosen. This hierarchy describes conditional actions within each frame and uses the model's existing text head. We route timing advantages to the token-group factor through event-causal masks derived from generated speech episodes, and route an LLM-judge semantic advantage to the conditional content factor. Across three Moshi seeds on Full-Duplex-Bench v1, HiPLEX reduces takeover rates during natural user pauses and backchannel opportunities, and shortens post-interruption response latency relative to GRPO, while maintaining comparable judged interruption-response quality. On Moshi and PersonaPlex, HiPLEX better matches pooled human turn-timing and backchannel-rate marginals than GRPO.
| Comments: | 34 pages, 9 figures, 17 tables, |
| Subjects: | Sound (cs.SD) |
| Cite as: | arXiv:2610.07727 [cs.SD] |
| (or arXiv:2610.07727v1 [cs.SD] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07727 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Kyudan Jung [view email]
[v1]
Tue, 6 Oct 2026 04:23:23 UTC (340 KB)
来源:HuggingFace Daily Papers · arxiv.org