arXiv:cs.AI· Kaiwen Luo, Chunxi Luo, Liang Lin, Yuxuan Li, Zhenhong Zhou, Junhao Dong, Yingjie Zhou, Zhendong Chu·· 4 小时前AI 评分34
EchoDistill:用噪声到干净自蒸馏提升大型音频语言模型鲁棒性
EchoDistill: Robust Large Audio Language Models via Noisy-to-Clean Self-Distillation
AI 导读
EchoDistill 是一种噪声到干净自蒸馏框架,在后期训练中以干净音频作为特权信息,让噪声输入的学生模型对齐干净条件下的语义,推理时仅保留学生模型、不增加额外开销。
正文
Abstract:Large Audio Language Models (LALMs) remain vulnerable to acoustic noise, which can obscure task-relevant evidence and produce unreliable responses. We propose EchoDistill, a noisy-to-clean self-distillation framework that uses clean audio as privileged information during post-training. A noisy-input student samples candidate responses reflecting its inference-time behavior, while a frozen copy of the same backbone processes the corresponding clean audio. EchoDistill combines masked response-token distillation, task-gated consistency shaping, and teacher-referenced group-relative optimization to align noisy-input generation with clean-conditioned semantics. Only the student is retained at inference time, introducing no additional inference cost. Across three LALM backbones and three audio domains at -10dB, EchoDistill improves average noisy-input accuracy by 1.63 percentage points over the strongest baseline. On Qwen2.5-Omni, it raises noisy-input accuracy from 59.33% to 62.94%, while clean-audio accuracy increases from 76.56% to 77.56%. Replacing matched audio with random, shuffled, or silent inputs reduces accuracy by 3.08-6.42 points, confirming that matched acoustic evidence contributes to its predictions. Additional evaluations show improvements on held-out additive noises and external benchmarks, while revealing that these gains do not reliably extend to non-additive distortions. These results demonstrate robust post-training improvements under severe additive noise without sacrificing clean-audio capability across diverse tasks.
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Sound (cs.SD) |
| Cite as: | arXiv:2605.23954 [cs.CL] |
| (or arXiv:2605.23954v2 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2605.23954 arXiv-issued DOI via DataCite |
Submission history
From: Kevin Luo [view email]
[v1]
Mon, 11 May 2026 06:30:25 UTC (7,840 KB)
[v2]
Fri, 2 Oct 2026 12:48:58 UTC (1,767 KB)
来源:arXiv:cs.AI · arxiv.org