arXiv:cs.AI· Shuran Ma, JiaLe Li, Yuxin Dong, Shan Zheng, Qingyun Jiang, Xiang Chen, Qi Zhu, Deyi Ji, Yifan Yang, Jianfeng Pan, Yu Tian, Xue Yang·· 4 小时前
超越视觉增强:AIMS 用自适应多源引导缓解 LVLM 幻觉
Beyond Visual Enhancement: Adaptive Multi-Context Steering to Mitigate LVLM Hallucinations
AI 导读
研究者提出 AIMS(Adaptive Information Multi-source Steering),一个轻量、免训练的框架,在解码时自适应协调视觉、预填文本与生成文本三类上下文。
正文
Authors:Shuran Ma, JiaLe Li, Yuxin Dong, Shan Zheng, Qingyun Jiang, Xiang Chen, Qi Zhu, Deyi Ji, Yifan Yang, Jianfeng Pan, Yu Tian, Xue Yang
Abstract:Hallucination remains a significant challenge in Large Vision-Language Models (LVLMs). Existing training-free methods generally mitigate hallucinations through contrastive decoding or visual enhancement, often increasing the relative influence of visual evidence during generation. This raises a fundamental question: Can LVLMs dynamically regulate the contributions of different context sources to suppress hallucinations? In this work, we investigate and quantify how LVLMs coordinate multiple context sources during decoding and examine how this intrinsic behavior can guide hallucination mitigation. We find that LVLMs exhibit an intrinsic vision-attending tendency that can guide adaptive visual steering, while textual contexts can also contribute to hallucination mitigation. Motivated by these findings, we propose AIMS (Adaptive Information Multi-source Steering), a lightweight training-free framework that adaptively coordinates visual, prefilled textual, and generated contexts during decoding. Specifically, AIMS constructs compact prototypes for the three context domains and estimates their affinities with the current query to determine head-wise steering weights. The resulting multi-source steering direction is applied to the query representation, enabling adaptive context integration without additional model training or auxiliary forward passes. Extensive experiments across multiple LVLMs and decoding strategies demonstrate that AIMS effectively mitigates object hallucination while maintaining competitive general-purpose multimodal capabilities.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.11907 [cs.CV] |
| (or arXiv:2610.11907v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11907 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Shuran Ma [view email]
[v1]
Thu, 8 Oct 2026 13:08:44 UTC (9,294 KB)
来源:arXiv:cs.AI · arxiv.org