跳到正文
arXiv:cs.AI· Jie Jin, Ziyin Ma, Min Yin, Jinyu Chen, Haigang Song, Zhikun Pang, Xiaowen Zhang·· 5 小时前AI 评分48

DuplexJev:冻结 LLM 单 Token 监督实现批量语音决策,无需解码

Batched Speech Decisions Without Decoding: Single-Token Supervision Lets a Frozen LLM Hear Beyond the Transcript

AI 导读

DuplexJev 将 ASR 编码器隐状态经小型连接器送入冻结 LLM,把每个问题读作选项上的单 Token 分布,无需解码即可完成决策,8-GPU 节点约 0.1 秒回答八段语音的 80 个决策。

正文

View PDF HTML (experimental)

Abstract:Full-duplex voice agents make many small, closed decisions, which current systems answer by slow autoregressive decoding. We propose DuplexJev, which feeds ASR-encoder hidden states through a small connector into a frozen LLM and reads each question as a single-token distribution over its options. Nothing is decoded, and an 8-GPU node answers 80 decisions about eight utterances in about 0.1 s. With a last-layer connector, spoken QA stays close to reading the transcript (90% vs. 91%). DuplexJev also hears the speaker: gender and emotion accuracy both reach 90% (from 55% and 28%) with a cross-attention connector, whose spoken QA drops by only 1 point (83% to 82%). We train decisions with cross-entropy on the read-out answer token, instead of the usual transcript distillation, whose teacher never hears the voice, and keep distillation for content. Encoders and LLMs are interchangeable; we release weights, training recipe, a batched-inference pipeline for full-duplex serving and a bilingual spoken-QA set.
Comments: 5 pages, 2 figures, 3 tables. Submitted to ICASSP 2027. Code and weights: this https URL
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.02638 [cs.AI]
  (or arXiv:2610.02638v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.02638

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Jinyu Chen [view email]
[v1] Fri, 2 Oct 2026 00:54:52 UTC (31 KB)

来源:arXiv:cs.AI · arxiv.org