arXiv:cs.LG· Lee Seung-woo, Bowen Qi·· 4 小时前AI 评分36
AdaLoop:面向音频语言模型的自适应深度潜在推理
AdaLoop: Adaptive-Depth Latent Reasoning for Audio Language Models
AI 导读
针对音频语言模型在细粒度声学分析任务上准确率骤降的问题,研究者提出轻量循环模块 AdaLoop,让模型按音频—问题对自行决定潜在精炼步数,并由学习到的停机机制在表示就绪后退出循环。
正文
Abstract:Large audio language models answer questions about speech, sound, and music, yet their accuracy drops sharply on tasks that need fine-grained acoustic analysis. Judging which of two speakers has the higher pitch demands iterative signal-level reasoning that a content question does not. Current models spend the same computational depth on both. We introduce AdaLoop, a lightweight recurrent module that learns how many latent refinement steps a given audio--question pair requires. A shared transformer block iterates over the audio representation, guided by the question, while a learned halting mechanism exits the loop once the representation is ready. AdaLoop adds fewer than 3\% of the base model's parameters and plugs into any audio encoder--language model pair without modifying either component. Evaluated on three architecturally distinct models across MMSU, MMAU-Pro, and MMAR, AdaLoop raises the average accuracy by 2.9 to 3.8 points, with the largest gains on perception-heavy subtasks where the model learns to apply deeper reasoning.
| Subjects: | Sound (cs.SD); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.06949 [cs.SD] |
| (or arXiv:2610.06949v1 [cs.SD] for this version) | |
| https://doi.org/10.48550/arXiv.2610.06949 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Bowen Qi [view email]
[v1]
Sat, 3 Oct 2026 11:27:41 UTC (25 KB)
来源:arXiv:cs.LG · arxiv.org