arXiv:cs.LG(机器学习,全量分类)· Pooneh Mousavi, Amir Ivry, Mirco Ravanelli, Cem Subakan·· 14 小时前AI 评分34
AnchorPrompt:面向鲁棒音频语言模型的自蒸馏软提示词
AnchorPrompt: Self-Distilled Soft Prompts for Robust Audio-Language Models
AI 导读
AnchorPrompt 是一种保持模型冻结的适配方法,在解码器输入端、音频与问题嵌入之间插入单个提示词向量块,并通过在多种音频与文本扰动上的自蒸馏训练这些向量。
正文
Abstract:Large audio-language models (LALMs) are sensitive to input perturbations, such as noise, waveform corruption, and adversarial injections. We propose AnchorPrompt, an efficient adaptation method that keeps the model frozen and learns a single block of prompt vectors inserted at the decoder input, between the audio and question embeddings. We train these vectors through self-distillation over diverse audio and text perturbations. To improve answer consistency and mitigate hallucination, we use the model's prediction on the clean recording as the target for answerable inputs, and assign a refusal target when the audio lacks sufficient evidence to answer. Furthermore, AnchorPrompt is perturbation-agnostic at inference, requiring no prior detection of perturbations and enabling zero-shot transfer to unseen distortions. We evaluate three LALMs across three benchmarks and show that AnchorPrompt improves answer consistency in most tested conditions. Clean accuracy improves in six of nine model-benchmark pairs, with minimal impact on the remainder of 1.2% at most. Crucially, AnchorPrompt reduces hallucinations under severe audio corruption while keeping false refusals on clean audio rare. Finally, these consistency gains transfer to unseen perturbations, such as choice permutations and reverberation.
| Subjects: | Sound (cs.SD); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.00706 [cs.SD] |
| (or arXiv:2610.00706v1 [cs.SD] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00706 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Pooneh Mousavi [view email]
[v1]
Wed, 30 Sep 2026 20:53:46 UTC (25 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org