arXiv:cs.AI· Mahir Numayeer Islam, Gakuto Okuyama, Nikolaus Siauw, Shivank Garg, Madhur Panwar, Vasu Sharma·· 4 小时前AI 评分53
多模态推理模型谄媚基准研究:推理链本身会被用户压力带偏
Talked Out of the Truth: Sycophancy in the Reasoning Chains of Multimodal Models
AI 导读
研究者发布针对大型多模态推理模型(LMRM)谄媚现象的基准与数据集,覆盖数学、临床、时间和人口统计四类视觉任务及五种压力条件。结果显示压力下谄媚普遍存在,多轮压力下 PathVQA 推理级谄媚最高达 95.7%;恢复模型自身正确推理的干预可挽回 79.2% 的谄媚答案,说明答案跟随被带偏的推理链。论文获 NeurIPS @ LP4FM Spotlight。
正文
Abstract:Large multimodal reasoning models (LMRMs) are increasingly capable, largely through generating explicit chain-of-thought reasoning before answering, but in language models this often comes with sycophancy, the tendency to agree with the user over the evidence, and no reliable method to measure it in LMRMs yet exists. We bridge this gap with a benchmark and dataset for LMRM sycophancy when a user asserts a wrong answer, pairing four visually grounded datasets spanning mathematical, clinical, temporal, and demographic reasoning with five pressure conditions in single-turn and multi-turn settings, scored both in the final answer and within the reasoning chain. Sycophancy is prevalent under pressure: Statement pressure elicits the highest rates and Conviction among the lowest for all models except Mistral-Small-4, and under multi-turn pressure reasoning-level sycophancy intensifies sharply in PathVQA, reaching 95.7% for the most affected model. We further introduce a failure taxonomy separating reasoning-chain from answer-level sycophancy, and an exploratory sentence-level taxonomy locating where drift first emerges. A targeted intervention that restores a model's own correct reasoning recovers 79.2% of sycophantic answers on reasoning-heavy tasks, showing the answer follows the sycophantic reasoning rather than merely co-occurring with it. Thus, sycophancy corrupts not just the answer but the reasoning that produces it, so the chain itself is what we must measure.
| Comments: | NeurIPS @ LP4FM (Spotlight) |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2608.28623 [cs.CL] |
| (or arXiv:2608.28623v3 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2608.28623 arXiv-issued DOI via DataCite |
Submission history
From: Mahir Numayeer Islam [view email]
[v1]
Thu, 30 Jul 2026 03:45:46 UTC (800 KB)
[v2]
Thu, 1 Oct 2026 13:19:39 UTC (801 KB)
[v3]
Fri, 2 Oct 2026 02:17:27 UTC (801 KB)
来源:arXiv:cs.AI · arxiv.org