跳到正文
arXiv:cs.LG· Aishik Nagar, Abhishek Vaidyanathan, Arun-Kumar Kaliya-Perumal, Elijah Tzen Hsuen Boey, Stefan Winkler·· 3 小时前AI 评分53

研究在 11 个开源 LLM 中发现可定位的临床概念中心

Clinical Concept Centers in LLMs

AI 导读

arXiv 论文(arXiv:2610.02829)在 11 个开源权重 LLM 的潜空间中发现可定位、可解释且因果驱动模型行为的临床概念中心。研究显示模型在对抗性角色 priming 下仍保持内部一致,沿这些中心做 steering 可带来下游临床性能提升,盲法临床医生验证表明概念中心的激活与使用能预测临床医生的偏好。

正文

View PDF HTML (experimental)

Abstract:Large language models are increasingly used in clinical settings. However, research into the reliability and performance of these models has focused almost entirely on the language substrate, scoring what the model says. Mechanistic interpretability has found that the latent space carries a higher fidelity of representation than the text: internal representations not only encode substantially more than the output verbalizes, but the stated reasoning also systematically omits features that causally drive the answer. An evaluation of model behavior in terms of mechanistic interpretability has not been explored in clinical decision support. In this work, we extend behavioral evaluation into the latent space and ask whether clinical concepts exist as locatable, causally used representations inside open-weight LLMs. We find dedicated clinical concept centers in the latent space of all eleven open models we test. These concept centers are interpretable, firing only on their aligned clinical narratives, and meaningfully and causally drive model behavior in both constrained and open-ended settings. They are not just analytical representations, but circuits that can be utilized in clinical practice, and we explore their use from the perspective of both evaluation and performance. From the evaluation standpoint, models stay internally coherent and keep using the relevant concept centers even under adversarial role-based priming, while aligned priming improves downstream clinical performance. From a performance perspective, we simulate realistic deployment settings and find that steering models along these centers leads to meaningful downstream improvements. Finally, we conduct a blinded clinician validation and find the activation and usage of these concept centers predicts clinicians preferences.
Subjects: Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as: arXiv:2610.02829 [cs.CL]
  (or arXiv:2610.02829v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.02829

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Stefan Winkler [view email]
[v1] Fri, 2 Oct 2026 05:22:12 UTC (689 KB)

来源:arXiv:cs.LG · arxiv.org