arXiv:cs.AI· Kaisong Zhang, Haotian Fang, Junmeng Zhou, Hang Lv, Yulan Pan, Yanchao Tan·· 3 小时前
CARing:用医学 token 与覆盖感知推理预测下次就诊诊断
Coverage-Aware Reasoning with Medical Tokens for Diagnosis Prediction
AI 导读
针对下次就诊诊断预测,研究者提出 CARing 框架,用残差量化将本体增强的疾病语义编码为紧凑的 Semantic IDs(SIDs),并通过覆盖奖励的强化学习与多阳性监督优化多标签覆盖。在 MIMIC-III 和 MIMIC-IV 上,CARing 的加权 F1 超过所有经 EHR 训练的基线,推理模式下 R@30 分别达 46.04% 和 46.52%,代码与日志已公开。
正文
Abstract:Large language models (LLMs) offer promising potential for next-visit diagnosis prediction, owing to their ability to integrate longitudinal clinical evidence and reason over it in natural language. However, reinforcement learning for LLM reasoning commonly rewards each trajectory according to the correctness of its final answer. In next-visit diagnosis prediction, multiple diagnoses can be simultaneously valid, but independently rewarding one diagnosis per trajectory does not distinguish repeated hits from coverage of different diagnoses. The policy can therefore concentrate on a few correct diagnoses, leaving others uncovered. Meanwhile, LLM tokenizers can split ICD codes into several generic tokens with limited clinical meaning, requiring multiple decoding steps to predict each diagnosis and hindering reasoning over a large disease vocabulary. To address both challenges, we propose CARing, a framework that represents diagnoses with compositional Semantic IDs (SIDs) and optimizes reasoning trajectories for multi-label coverage. Concretely, we first encode ontology-enriched disease semantics into compact SIDs through residual quantization, and ground the resulting SID tokens in natural language and longitudinal EHR contexts through multi-task alignment and reasoning-enriched training to unlock transferable LLM reasoning. CARing further improves unordered multi-label prediction through a coverage reward for reinforcement learning and multi-positive supervision. At inference time, the model supports both efficient direct constrained decoding and multi-chain reasoning with rank fusion. On MIMIC-III and MIMIC-IV, CARing exceeds all EHR-trained baselines in weighted F1 and attains the highest top-k recall at every reported cutoff, including R@30 of 46.04% and 46.52% in reasoning mode. Our codes and logs are available at this https URL.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.10641 [cs.LG] |
| (or arXiv:2610.10641v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10641 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Kaisong Zhang [view email]
[v1]
Wed, 7 Oct 2026 14:39:18 UTC (426 KB)
来源:arXiv:cs.AI · arxiv.org