arXiv:cs.AI· Min Zeng, Rui Zhang·· 5 小时前AI 评分45
研究发现 LLM 临床判断随患者证据演变时更新不可靠
Large language models exhibit unreliable updating of clinical judgment as patient evidence evolves
AI 导读
一项基于电子健康记录中匹配的 ICU 病程轨迹的研究发现,多种 LLM 在患者证据演变时对临床判断的更新不可靠:当估计值发生变化时,以先前判断为条件反而更常增大而非降低预测误差。受控干预揭示两种失效模式——模型对恶化证据的反应强于同等程度的改善证据,且当前证据固定时,将先验风险从 10% 提高到 90% 会使估计值偏移 26.2 个百分点;提示词无法恢复可靠更新。
正文
Abstract:Large language models (LLMs) are increasingly explored for clinical reasoning, but whether they appropriately revise judgments as patient evidence evolves remains unclear. We evaluated longitudinal belief updating using matched intensive-care trajectories from electronic health records. Across diverse LLMs, conditioning on a preceding judgment more often increased than reduced prediction error when estimates changed, replicated for a second endpoint. Controlled interventions revealed two failure modes. First, with preceding assessment fixed, models responded more strongly to worsening than matched improving respiratory evidence; this asymmetry persisted after headroom normalization at moderate and strong evidence levels. Second, with current evidence fixed, increasing prior risk from 10% to 90% shifted estimates by 26.2 percentage points, demonstrating causal influence of prior model beliefs. Prompting did not restore reliable updating. Evidence-Validated Longitudinal Update (EVLU) identified fewer, more reliable revisions, revealing a reliability-coverage trade-off. These findings establish longitudinal belief updating as a distinct dimension of LLM reliability.
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.02684 [cs.AI] |
| (or arXiv:2610.02684v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02684 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Min Zeng [view email]
[v1]
Fri, 2 Oct 2026 02:05:01 UTC (488 KB)
来源:arXiv:cs.AI · arxiv.org