跳到正文
arXiv:cs.AI· Min Zeng, Rui Zhang·· 5 小时前AI 评分45

研究发现 LLM 临床判断随患者证据演变时更新不可靠

Large language models exhibit unreliable updating of clinical judgment as patient evidence evolves

AI 导读

一项基于电子健康记录中匹配的 ICU 病程轨迹的研究发现,多种 LLM 在患者证据演变时对临床判断的更新不可靠:当估计值发生变化时,以先前判断为条件反而更常增大而非降低预测误差。受控干预揭示两种失效模式——模型对恶化证据的反应强于同等程度的改善证据,且当前证据固定时,将先验风险从 10% 提高到 90% 会使估计值偏移 26.2 个百分点;提示词无法恢复可靠更新。

正文

View PDF HTML (experimental)

Abstract:Large language models (LLMs) are increasingly explored for clinical reasoning, but whether they appropriately revise judgments as patient evidence evolves remains unclear. We evaluated longitudinal belief updating using matched intensive-care trajectories from electronic health records. Across diverse LLMs, conditioning on a preceding judgment more often increased than reduced prediction error when estimates changed, replicated for a second endpoint. Controlled interventions revealed two failure modes. First, with preceding assessment fixed, models responded more strongly to worsening than matched improving respiratory evidence; this asymmetry persisted after headroom normalization at moderate and strong evidence levels. Second, with current evidence fixed, increasing prior risk from 10% to 90% shifted estimates by 26.2 percentage points, demonstrating causal influence of prior model beliefs. Prompting did not restore reliable updating. Evidence-Validated Longitudinal Update (EVLU) identified fewer, more reliable revisions, revealing a reliability-coverage trade-off. These findings establish longitudinal belief updating as a distinct dimension of LLM reliability.
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as: arXiv:2610.02684 [cs.AI]
  (or arXiv:2610.02684v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.02684

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Min Zeng [view email]
[v1] Fri, 2 Oct 2026 02:05:01 UTC (488 KB)

来源:arXiv:cs.AI · arxiv.org