arXiv:cs.AI· Shrihari Dumbre, Bikash Santra·· 10 小时前AI 评分32
多模态知识蒸馏实现胃腺癌全切片图像分类
Multimodal Knowledge Distillation for Gastric Adenocarcinoma Classification from Whole-Slide Images
AI 导读
研究提出多模态知识蒸馏(MKD)框架,用低秩多模态融合(LMF)在训练中结合预训练 WSI 图像编码器与临床文本编码器,让教师模型学习图像-文本融合表征,学生模型蒸馏后仅凭图像即可推理。该方法在 PatchGastric 基准上平均准确率比 SOTA 至少高 3.35%,且不依赖 Transformer 融合、多任务学习或大语言模型。源码已公开。
正文
Abstract:Gastric adenocarcinoma (GA) is a leading cause of cancer-related mortality worldwide, and accurate histopathological subtype classification from whole-slide images (WSIs) is essential for effective treatment planning. While multimodal approaches that integrate pathology report text with WSIs can improve classification, existing methods often depend on computationally expensive transformer architectures and large language models. We propose a multimodal knowledge distillation (MKD) framework that combines a pretrained WSI image encoder and a clinical text encoder using Low-Rank Multimodal Fusion (LMF) to efficiently model cross-modal interactions during training. Each WSI is represented as a bag of patches paired with a slide-level diagnostic caption. The teacher model learns fused image-text representations for subtype classification, while the student model distills this knowledge to enable accurate image-only inference. We evaluate our method on the PatchGastric benchmark dataset and achieve at least 3.35% higher mean accuracy than state-of-the-art approaches, without relying on transformer-based fusion, multi-task learning, or large language models. The source code is available at this https URL.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.07913 [cs.CV] |
| (or arXiv:2610.07913v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07913 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Bikash Santra [view email]
[v1]
Tue, 6 Oct 2026 07:56:54 UTC (5,469 KB)
来源:arXiv:cs.AI · arxiv.org