arXiv:cs.LG(机器学习,全量分类)· Yuxiang Wang, Kunyu Feng, Yuancheng Wang, Zihang Liu, Shengbo Cai, Qinke Ni, Wan Lin, Tao Feng, Yingda shen, Ming-Hao Hsu, Zhixian Zhao, Liqiang Zhang, Teddy Sun, Steve Yves, Zhizheng Wu·· 17 小时前AI 评分42
AURAL:面向语音语言模型的自适应潜在推理与联合分块方法
AURAL: Adaptive Latent Reasoning with Joint Chunk for Speech Language Models
AI 导读
研究者提出 AURAL,通过在潜在空间建模多条推理路径并联合预测未来状态分块,减少顺序前向计算与推理延迟。团队构建了 AuralReason-683K 数据集,包含 68.3 万条双语语音语句(约 1000 小时),覆盖情感识别、共情对话与通用推理。AURAL-RL 在 Qwen2.5-Omni 上将首个答案 token 的时间从 1.22 秒降至 0.10 秒,提速 11.8 倍。
正文
Authors:Yuxiang Wang, Kunyu Feng, Yuancheng Wang, Zihang Liu, Shengbo Cai, Qinke Ni, Wan Lin, Tao Feng, Yingda shen, Ming-Hao Hsu, Zhixian Zhao, Liqiang Zhang, Teddy Sun, Steve Yves, Zhizheng Wu
Abstract:Model intelligence and fast response jointly shape the quality of interaction with speech language models, yet remain difficult to achieve together. Explicit chain-of-thought (CoT) improves reasoning and audio understanding, but generating intermediate reasoning tokens delays responses. Describing fine-grained acoustic cues further lengthens CoT and increases latency. Latent reasoning can reduce this overhead, yet existing methods often trail CoT and remain limited by single-path supervision and reasoning budgets that do not adapt to problem difficulty. We introduce AURAL, which models a distribution over multiple plausible reasoning continuations in latent space and jointly predicts chunks of future states to reduce sequential forward passes and reasoning latency. To provide initial supervision for latent reasoning, we construct AuralReason-683K: 683K bilingual speech utterances (about 1,000 hours) with concise CoT for emotion recognition, empathetic dialogue, and general reasoning. AURAL-RL then explores beyond these traces, rewarding concise reasoning that yields high-quality answers and adapting reasoning effort to each problem. Across two backbones, AURAL-RL achieves performance comparable to CoT-RL, with larger gains over the respective supervised checkpoints on most metrics. Analysis further shows that harder questions elicit more latent reasoning steps. On Qwen2.5-Omni, it reduces time to the first answer token by 11.8x, from 1.22 to 0.10 s, versus 0.05 s for direct answering.
| Subjects: | Computation and Language (cs.CL); Machine Learning (cs.LG); Sound (cs.SD) |
| Cite as: | arXiv:2610.01560 [cs.CL] |
| (or arXiv:2610.01560v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.01560 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yuxiang Wang [view email]
[v1]
Thu, 1 Oct 2026 12:28:03 UTC (1,317 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org