跳到正文
arXiv:cs.AI· Kaiwen Luo, Chunxi Luo, Liang Lin, Yuxuan Li, Zhenhong Zhou, Junhao Dong, Yingjie Zhou, Zhendong Chu·· 4 小时前AI 评分34

EchoDistill:用噪声到干净自蒸馏提升大型音频语言模型鲁棒性

EchoDistill: Robust Large Audio Language Models via Noisy-to-Clean Self-Distillation

AI 导读

EchoDistill 是一种噪声到干净自蒸馏框架,在后期训练中以干净音频作为特权信息,让噪声输入的学生模型对齐干净条件下的语义,推理时仅保留学生模型、不增加额外开销。

正文

View PDF HTML (experimental)

Abstract:Large Audio Language Models (LALMs) remain vulnerable to acoustic noise, which can obscure task-relevant evidence and produce unreliable responses. We propose EchoDistill, a noisy-to-clean self-distillation framework that uses clean audio as privileged information during post-training. A noisy-input student samples candidate responses reflecting its inference-time behavior, while a frozen copy of the same backbone processes the corresponding clean audio. EchoDistill combines masked response-token distillation, task-gated consistency shaping, and teacher-referenced group-relative optimization to align noisy-input generation with clean-conditioned semantics. Only the student is retained at inference time, introducing no additional inference cost. Across three LALM backbones and three audio domains at -10dB, EchoDistill improves average noisy-input accuracy by 1.63 percentage points over the strongest baseline. On Qwen2.5-Omni, it raises noisy-input accuracy from 59.33% to 62.94%, while clean-audio accuracy increases from 76.56% to 77.56%. Replacing matched audio with random, shuffled, or silent inputs reduces accuracy by 3.08-6.42 points, confirming that matched acoustic evidence contributes to its predictions. Additional evaluations show improvements on held-out additive noises and external benchmarks, while revealing that these gains do not reliably extend to non-additive distortions. These results demonstrate robust post-training improvements under severe additive noise without sacrificing clean-audio capability across diverse tasks.
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Sound (cs.SD)
Cite as: arXiv:2605.23954 [cs.CL]
  (or arXiv:2605.23954v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2605.23954

arXiv-issued DOI via DataCite

Submission history

From: Kevin Luo [view email]
[v1] Mon, 11 May 2026 06:30:25 UTC (7,840 KB)
[v2] Fri, 2 Oct 2026 12:48:58 UTC (1,767 KB)

来源:arXiv:cs.AI · arxiv.org