arXiv:cs.AI· Maryam Baizhigitova, Andrew Seohwan Yu, Po-Hao Chen, Naveen Subhas, Sixu Chen, Xinxin Wang, Kunio Nakamura, Richard Lartey, Xiaojuan Li, Mingrui Yang·· 6 小时前AI 评分33
Knee3DVLM:面向膝关节 MRI 全面评估的双序列全容积视觉语言建模
Knee3DVLM: Dual-Sequence Full-Volume Vision-Language Modeling for Comprehensive Knee MRI Assessment
AI 导读
Knee3DVLM 是一个序列感知的视觉语言模型,利用全容积 DESS 与液体敏感 TSE MRI 预测 57 个源自 MOAKS 的解剖定位二分类诊断目标用于结构化报告。
正文
Abstract:Vision-language models (VLMs) are increasingly being applied to three-dimensional medical imaging, but their application to knee MRI remains limited, particularly for interpreting the complementary sequences used in clinical practice. We introduce Knee3DVLM, a sequence-aware VLM that uses full-volume DESS and fluid-sensitive TSE MRI to predict 57 anatomically resolved binary diagnostic targets derived from the MRI Osteoarthritis Knee Score (MOAKS) for structured reporting. We evaluated DESS-only, TSE-only, and paired DESS-TSE configurations using subject-disjoint Osteoarthritis Initiative partitions. In a held-out cohort of 1,074 examinations, the fused model achieved 72.98% average accuracy, 71.17% balanced accuracy, 78.96% mean ROC-AUC, and 78.74% macro ROC-AUC, the highest values among the three configurations. In a secondary multiclass analysis aligned with the released 3DReasonKnee cohort, Knee3DVLM was numerically higher than the strongest reported 3DReasonKnee configuration across five pathology categories. These findings support dual-sequence full-volume modeling for comprehensive knee MRI assessment.
| Comments: | 11 pages, 2 figures, 5 tables |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.08482 [cs.CV] |
| (or arXiv:2610.08482v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08482 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Maryam Baizhigitova [view email]
[v1]
Tue, 6 Oct 2026 14:58:53 UTC (895 KB)
来源:arXiv:cs.AI · arxiv.org