跳到正文
arXiv:cs.LG· Taha Emre, Arunava Chakravarty, Thomas Pinetz, Dmitrii Lachinov, Martin J. Menten, Hendrik Scholl, Sobha Sivaprasad, Daniel Rueckert, Andrew Lotery, Stefan Sacu, Ursula Schmidt-Erfurth, Hrvoje Bogunovi\'c·· 7 小时前AI 评分34

STAMP:面向纵向医学图像的随机 Siamese MAE 预训练框架

Stochastic Siamese MAE Pretraining for Longitudinal Medical Images

AI 导读

研究者提出 STAMP,一种 Siamese MAE 框架,通过随机过程对两个输入体积之间的时间差进行条件编码,将 MAE 重建损失重新表述为条件变分推理目标,从而以随机方式学习时间动态。

正文

Authors:Taha Emre, Arunava Chakravarty, Thomas Pinetz, Dmitrii Lachinov, Martin J. Menten, Hendrik Scholl, Sobha Sivaprasad, Daniel Rueckert, Andrew Lotery, Stefan Sacu, Ursula Schmidt-Erfurth, Hrvoje Bogunović

View PDF HTML (experimental)

Abstract:Temporally aware image representations are crucial for capturing disease progression in 3D volumes of longitudinal medical datasets. However, recent state-of-the-art self-supervised learning approaches like Masked Autoencoding (MAE), despite their strong representation learning capabilities, lack temporal awareness. In this paper, we propose STAMP (Stochastic Temporal Autoencoder with Masked Pretraining), a Siamese MAE framework that encodes temporal information through a stochastic process by conditioning on the time difference between the 2 input volumes. Unlike deterministic Siamese approaches, which compare scans from different time points but fail to account for the inherent uncertainty in disease evolution, STAMP learns temporal dynamics stochastically by reframing the MAE reconstruction loss as a conditional variational inference objective. We evaluated STAMP on two OCT and one MRI datasets with multiple visits per patient. STAMP pretrained ViT models outperformed both existing temporal MAE methods and foundation models on different late stage Age-Related Macular Degeneration and Alzheimer's Disease progression prediction which require models to learn the underlying non-deterministic temporal dynamics of the diseases.
Comments: Provisional Accept at IEEE TMI. Code is available in this https URL
Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
Cite as: arXiv:2512.23441 [cs.LG]
  (or arXiv:2512.23441v2 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2512.23441

arXiv-issued DOI via DataCite

Submission history

From: Taha Emre [view email]
[v1] Mon, 29 Dec 2025 13:00:12 UTC (6,585 KB)
[v2] Tue, 6 Oct 2026 13:10:26 UTC (6,645 KB)

来源:arXiv:cs.LG · arxiv.org