跳到正文
arXiv:cs.AI· Kaiser Sun, Bernal Jimenez Gutierrez, Hongjun Liu, Jingyu Zhang, Jie Gao, Mark Dredze, Daniel Khashabi·· 3 小时前

准确但不高傲:评估知识冲突下 LLM 智能体的认知谦逊

Accurate but Not Humble: Evaluating Epistemic Humility in LLM Agents under Knowledge Conflict

AI 导读

研究者提出用"认知谦逊"(EH)评估 LLM 智能体在知识冲突下的表现,通过 Identify、Solve、Escalate(ISE)三个轨迹级行为维度,在受控冲突和多步执行中的自然冲突两种场景下测试了四个智能体。

正文

View PDF HTML (experimental)

Abstract:When retrieved evidence contradicts an agent's prior beliefs, does it revise its answer, acknowledge uncertainty, or persist with an incorrect conclusion? Existing evaluations of agentic systems focus primarily on task success, offering limited insight into how agents handle such conflicts. We propose to evaluate agents on epistemic humility (EH): the agent's willingness to recognize, act on, and communicate uncertainty during task execution. We operationalize EH through three trajectory-level behavioral dimensions: Identify, Solve, and Escalate (ISE). Through knowledge conflict, situations where the backbone language model's parametric knowledge contradicts the evidence it encounters, or where two contextual sources disagree, we evaluate two conflict settings: (1) controlled conflict and (2) naturally occurring conflict during multi-step agentic execution, each paired with matched no-conflict controls. Evaluating four agents, we find that higher task accuracy does not necessarily correspond to greater epistemic humility: some high-accuracy configurations recognize conflicts during execution but do not communicate unresolved uncertainty in their incorrect final answers. Trajectory-level analysis further reveals that agents frequently detect conflicts in early steps of execution but fail to maintain or resolve them in later steps. Finally, we show that model-level interventions can improve EH, but often at the cost of task accuracy, suggesting that epistemic humility emerges from the interaction among the backbone model, agent harness, and evaluation environment.
Comments: EMNLP 2026 Camera Ready
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as: arXiv:2610.12360 [cs.AI]
  (or arXiv:2610.12360v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.12360

arXiv-issued DOI via DataCite (pending registration)

Journal reference: EMNLP2026

Submission history

From: Kaiser Sun [view email]
[v1] Thu, 8 Oct 2026 17:25:06 UTC (790 KB)

来源:arXiv:cs.AI · arxiv.org