跳到正文
arXiv:cs.AI· Maryam Baizhigitova, Andrew Seohwan Yu, Po-Hao Chen, Naveen Subhas, Sixu Chen, Xinxin Wang, Kunio Nakamura, Richard Lartey, Xiaojuan Li, Mingrui Yang·· 6 小时前AI 评分33

Knee3DVLM:面向膝关节 MRI 全面评估的双序列全容积视觉语言建模

Knee3DVLM: Dual-Sequence Full-Volume Vision-Language Modeling for Comprehensive Knee MRI Assessment

AI 导读

Knee3DVLM 是一个序列感知的视觉语言模型,利用全容积 DESS 与液体敏感 TSE MRI 预测 57 个源自 MOAKS 的解剖定位二分类诊断目标用于结构化报告。

正文

View PDF HTML (experimental)

Abstract:Vision-language models (VLMs) are increasingly being applied to three-dimensional medical imaging, but their application to knee MRI remains limited, particularly for interpreting the complementary sequences used in clinical practice. We introduce Knee3DVLM, a sequence-aware VLM that uses full-volume DESS and fluid-sensitive TSE MRI to predict 57 anatomically resolved binary diagnostic targets derived from the MRI Osteoarthritis Knee Score (MOAKS) for structured reporting. We evaluated DESS-only, TSE-only, and paired DESS-TSE configurations using subject-disjoint Osteoarthritis Initiative partitions. In a held-out cohort of 1,074 examinations, the fused model achieved 72.98% average accuracy, 71.17% balanced accuracy, 78.96% mean ROC-AUC, and 78.74% macro ROC-AUC, the highest values among the three configurations. In a secondary multiclass analysis aligned with the released 3DReasonKnee cohort, Knee3DVLM was numerically higher than the strongest reported 3DReasonKnee configuration across five pathology categories. These findings support dual-sequence full-volume modeling for comprehensive knee MRI assessment.
Comments: 11 pages, 2 figures, 5 tables
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.08482 [cs.CV]
  (or arXiv:2610.08482v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2610.08482

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Maryam Baizhigitova [view email]
[v1] Tue, 6 Oct 2026 14:58:53 UTC (895 KB)

来源:arXiv:cs.AI · arxiv.org