跳到正文
arXiv:cs.AI· Hyunjae Ra, Aecheon Jung, Jungin Park, Sungeun Hong·· 5 小时前AI 评分37

RED:面向音视频幻觉缓解的相关证据解码方法

Relevant Evidence Decoding for Audio-Visual Hallucination Mitigation

AI 导读

针对音视频大语言模型(AV-LLMs)的跨模态幻觉问题,研究者提出免训练方法 Relevant Evidence Decoding(RED),用点互信息量化音频与视频对预测的贡献,并按问题所需证据类型选择性增强。

正文

View PDF HTML (experimental)

Abstract:Audio-Visual Large Language Models (AV-LLMs) remain prone to cross-modal hallucinations, where one modality incorrectly affects predictions about another. Although contrastive decoding reduces hallucinations in vision-language models, its direct extension to AV-LLMs overlooks a key challenge: different questions require different perceptual evidence, including audio, video, or their interaction. Notably, we observe that joint audio-visual inference can weaken the prediction even when a model can recover the correct answer from a single informative modality. For example, when asked which instrument is heard, a model may correctly predict violin from the audio alone. Once a video showing a guitar is added, its confidence in violin may drop. In this paper, we introduce Relevant Evidence Decoding (RED), a training-free method that identifies question-relevant evidence and selectively strengthens its contribution. RED uses pointwise mutual information to quantify the predictive support provided by audio and video beyond the question alone. It decomposes their joint contribution into audio, video, and residual interaction components. A question-only inference pass determines the required evidence type, after which the model augments the original audio-visual prediction with the corresponding PMI contribution. Across three audio-visual hallucination benchmarks and three AV-LLMs, RED improves accuracy over standard decoding by up to 7% on CMM, 6.3% on AVHBench, and 3.8% on SVHalluc, with an average relative time to first token of 1.5x standard decoding.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.02976 [cs.AI]
  (or arXiv:2610.02976v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.02976

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Hyunjae Ra [view email]
[v1] Fri, 2 Oct 2026 08:09:19 UTC (2,817 KB)

来源:arXiv:cs.AI · arxiv.org