arXiv:cs.LG· Sheir A. Zaheer, Jihwan Moon, Chan Y. Park·· 4 小时前AI 评分29
REViT-v2:面向等变特征提取的层级窗口化旋转-反射等变 ViT
REViT-v2: Hierarchical Windowed Roto-reflection Equivariant ViT for Equivariant Feature Extraction
AI 导读
REViT-v2 是一种可扩展的旋转-反射群等变视觉 Transformer,基于窗口化群卷积自注意力与层级特征架构,可扩展到百万级参数规模并在 ImageNet 这类实用尺寸图像的大数据集上训练。其代码与预训练权重已公开,论文被 NeurIPS NeurREPS workshop 2026 接收。
正文
Abstract:We propose a scalable roto-reflection-group-equivariant vision transformer based on windowed group-convolutional self-attention and a hierarchical feature architecture. We demonstrate that our approach can be scaled to group-equivariant vision transformers (ViTs) with millions of parameters and large datasets with practically sized images, i.e., ImageNet. The code and pretrained weights for the proposed Hierarchical Windowed Roto-reflection Equivariant ViTs (REViT-v2) are available at this https URL.
| Comments: | 7 pages, Accepted for presentation at NeurIPS NeurREPS workshop 2026 |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.07585 [cs.CV] |
| (or arXiv:2610.07585v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07585 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Sheir A. Zaheer [view email]
[v1]
Tue, 6 Oct 2026 01:23:15 UTC (20 KB)
来源:arXiv:cs.LG · arxiv.org