跳到正文
arXiv:cs.AI· Asim Khan, Samee Ullah Khan, Dwarikanath Mahapatra·· 6 小时前AI 评分34

MedCORE:基于临床标准的可解释医学影像诊断推理框架

MedCORE: Criteria-Grounded Clinical Reasoning for Interpretable Medical Image Diagnosis

AI 导读

MedCORE 是一个将临床诊断标准显式嵌入视觉-语言架构的结构化诊断框架,把诊断过程分解为临床定义的各项标准并逐一定位到图像区域。它在 ISIC 2018 皮肤镜分类上达到 89.2% 准确率、85.7% macro-F1 和 96.4% AUC,在 BUSI 乳腺超声上达到 96.1% 准确率、95.2% macro-F1 和 98.4% AUC。

正文

View PDF HTML (experimental)

Abstract:Clinical diagnosis is inherently a structured reasoning process, yet existing deep learning models often bypass this structure by mapping image features directly to disease labels without explicitly interrogating the morphological and textural criteria that clinicians systematically evaluate. This limits diagnostic transparency and may compromise safe clinical deployment. We present MedCORE (Medical Criteria-Oriented Reasoning and Evidence), a structured diagnostic framework that operationalizes clinical reasoning within a vision-language architecture. For each input image, MedCORE decomposes the diagnostic process into clinically defined criteria, spatially localizes each criterion to diagnostically relevant image regions, encodes evidence through multi-scale representations that capture macro-structural and micro-textural pathological characteristics, and refines criterion representations using a Graph Attention Network that explicitly models inter-criteria dependencies. Criterion representations are further aligned with clinical text descriptors, reinforced through class-wise visual prototypes, and aggregated using uncertainty-calibrated weighting that proportionally discounts low-confidence diagnostic evidence. MedCORE is validated across three clinically heterogeneous imaging modalities, including dermoscopic lesion classification on ISIC 2018, breast ultrasound lesion characterization on BUSI, and diabetic retinopathy grading on IDRiD. Quantitatively, MedCORE achieves 89.2% accuracy, 85.7% macro-F1, and 96.4% AUC on ISIC 2018; 96.1% accuracy, 95.2% macro-F1, and 98.4% AUC on BUSI; and 84.3% accuracy, 80.2% macro-F1, and 92.8% AUC on IDRiD. These results demonstrate consistent improvements over strong CNN, transformer, biomedical vision-language, concept-based, and prototype-based baselines.
Comments: 16 pages, 4 figures, conference
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.08528 [cs.CV]
  (or arXiv:2610.08528v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2610.08528

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Asim Khan [view email]
[v1] Tue, 6 Oct 2026 15:23:01 UTC (4,280 KB)

来源:arXiv:cs.AI · arxiv.org