arXiv:cs.LG· Eduardo Santos-Escriche, Valerie Engelmayer, Ya-Wei Eileen Lin, Stefanie Jegelka·· 4 小时前AI 评分34
Transformer 如何学习表征对称性?
How Do Transformers Learn to Represent Symmetries?
AI 导读
研究者用有限数据增强训练 vanilla Transformer,发现不同对称群的习得难度存在明确排序:非保角对称性最易学,其次是保角对称性,平移、旋转、缩放等基础保角子群最难。针对基础保角群,他们还分析了模型的泛化行为与内部结构,识别出可解释的近似不变性机制,并证明这些机制可进一步支撑等变函数的学习。该工作已被 NeurIPS 2026 接收。
正文
Abstract:Training Transformer-based architectures with finite data augmentation has become an increasingly popular approach in geometric machine learning. Despite its empirical success, the interplay between the Transformer architecture, invariance to different symmetries, and augmentation budgets remains underexplored. In this paper, we study the ability of a vanilla Transformer to learn various symmetries through finite data augmentation for point cloud datasets. We identify an ordering of increasing learnability across the following symmetry groups: (i) non-angle-preserving symmetries, (ii) angle-preserving symmetries, and (iii) base angle-preserving subgroups, such as translation, rotation, and scale. For the base angle-preserving groups, we further investigate the Transformer's extrapolation behavior and conduct a structural analysis of the trained models, allowing us to identify interpretable mechanisms that induce invariance. Finally, we extend our analysis to equivariant functions and show that the detected mechanisms for approximate invariance can also provide a key building block for learned equivariance. Our project page is available at this https URL
| Comments: | Accepted at NeurIPS 2026 |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.10305 [cs.LG] |
| (or arXiv:2610.10305v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10305 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Eduardo Santos Escriche [view email]
[v1]
Wed, 7 Oct 2026 16:01:51 UTC (4,542 KB)
来源:arXiv:cs.LG · arxiv.org