跳到正文
arXiv:cs.CL· Jongwook Yoon, Jongwon Lim, Sungjib Lim, Woojin Cho, Yohan Jo·· 3 小时前AI 评分43

LLM 如何处理否定?机制分析揭示否定失败根源

How Do LLMs Change Predictions Under Negation?

AI 导读

一项 arXiv 研究评估了近期开源与闭源 LLM 在否定基准上的表现,发现它们在 37-71% 的情况下对否定问句重复原答案,例如被问"什么是西班牙非首都"时仍答"Madrid"。

正文

View PDF HTML (experimental)

Abstract:Negation is an essential feature of human language, yet large language models (LLMs) remain unreliable in processing it. We evaluate recent open-source and closed-source LLMs on our negation benchmark and find that, in 37-71% of cases, they repeat the same answer under negation (e.g., "Madrid" for "What is not the capital of Spain?"). To understand and address this brittleness, we mechanistically examine how models operate under negation. Our main finding is that specialized attention heads and MLP neurons jointly implement negation by (1) suppressing retrieval of the original answer (e.g., "Madrid") while (2) promoting a favored candidate within the answer category (e.g., "Paris"). This contrasts with accounts of human negation processing, in which information about the original answer helps to determine what should be excluded. Furthermore, we find that this difference from human processing is a key source of negation failures: the model's mechanism relies on suppressing the original answer rather than using it to determine what to exclude, so the model can repeat the original answer when suppression is too weak or when a bias toward particular answers prevents it from selecting an alternative. To address this weakness in the model's negation mechanism, we propose a training objective that requires larger shifts in answer preference for more confident original predictions, and show that it reduces negation failures with less degradation of general capabilities than standard fine-tuning baselines. Together, our results demonstrate how mechanistic analysis can reveal why a linguistic capability fails and guide training that targets the underlying limitation.
Comments: Under Review
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.09571 [cs.CL]
  (or arXiv:2610.09571v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.09571

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Jongwon Lim [view email]
[v1] Wed, 7 Oct 2026 07:12:16 UTC (1,480 KB)

来源:arXiv:cs.CL · arxiv.org