arXiv:cs.LG· AmirEhsan Khorashadizadeh, Benjam\'in B\'ejar·· 5 小时前AI 评分43
TomoTransformer:面向 CT 重建的基础模型
TomoTransformer: Towards a Foundation Model for CT Reconstruction
AI 导读
TomoTransformer 是一个基于 Transformer 的 CT 重建基础模型,将每个局部滤波投影视为独立 token,通过自注意力预测缺失视角,无需重训练即可处理任意数量、任意角度和任意探测器尺寸的输入投影。该模型在多个稀疏视角基准数据集上显著优于 ViewTrans 等多用途模型,并匹配或超越协议专用基线,还在 X 射线同步加速器采集的纳米级脑数据上展现出稳健的零样本泛化能力。
正文
Abstract:Supervised deep learning has advanced sparse-view tomographic reconstruction. However, conventional models, which typically map filtered back-projection (FBP) images or sinograms to clean reconstructions, are brittle under distribution shifts. Because they require retraining whenever projection counts and angles, detector resolutions, or data distributions change, their deployment in real-world applications remains limited. To address this, we introduce TomoTransformer, a transformer-based architecture that treats each \textit{local} filtered projection as an individual token and predicts missing views via self-attention. Crucially, TomoTransformer operates in a \emph{back-projection space} that separates projections across spatial locations, making view interpolation geometrically well-posed and invariant to detector size. This design yields a single foundation model that can process any number of input projections, at arbitrary angular locations and detector dimensions, and query any number of target angles without retraining. Trained on a large-scale dataset spanning diverse medical CT anatomies and natural images, TomoTransformer generalizes effectively across anatomies, materials, and resolutions. Extensive evaluations on several benchmark sparse-view datasets show that TomoTransformer significantly outperforms concurrent multi-purpose models like ViewTrans and matches or exceeds strong protocol-specific baselines, while remaining fully agnostic to the number of input and target projections. Furthermore, the model demonstrates robust zero-shot generalization on real experimental nanoscale brain data collected from an X-ray synchrotron, showcasing its practical utility for real-world applications.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG) |
| Cite as: | arXiv:2609.37605 [cs.CV] |
| (or arXiv:2609.37605v2 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2609.37605 arXiv-issued DOI via DataCite |
Submission history
From: AmirEhsan Khorashadizadeh [view email]
[v1]
Tue, 29 Sep 2026 13:53:27 UTC (13,527 KB)
[v2]
Fri, 2 Oct 2026 11:56:50 UTC (13,527 KB)
来源:arXiv:cs.LG · arxiv.org