arXiv:cs.LG· Raneem Mahajne, Toviah Moldwin·· 4 小时前AI 评分35
完全可解释的极简 Transformer:从几何到算法
Fully Interpretable Minimal Transformers: From Geometry to Algorithm
AI 导读
研究者提出将 Transformer 的嵌入维度和注意力头大小都限制为 2,从而实现对内部表示的全二维可视化,并训练其在数字序列中识别 '+' 后输出最近出现的偶数。训练后,嵌入、Q/K/V 变换、注意力矩阵与决策边界均可直接观察,模型学到的几何结构可被解读为逐步算法。该工作提供了 15 张图表的可解释性可视化套件与训练动态动画。
正文
Abstract:We present a framework for building and interpreting minimal transformer models. By constraining a transformer's embedding dimension and head size to 2, we enable full two-dimensional visualization of its internal representations. Embeddings, query/key/value transforms, attention outputs, residual streams, and decision boundaries can all be seen directly. Our central claim is that the learned geometry implies an algorithm; the arrangement of points and boundaries in R^2 can be read as a step-by-step procedure. We train a transformer on a simple task where it must produce the most recently observed even number whenever the '+' operator appears in a sequence of digits. Once trained, we visually walk through every step of the transformer's computation. We show how the model embeds the tokens and their respective positions in the sequence, transforms them via the Q, K, and V matrices, uses the dot product between the Q and K representations to form the attention matrix, and uses the attention matrix to select values that move the representation of each input token to the region of the domain of the output layer that will correctly predict the next token. We introduce a suite of interpretability visualizations that make the algorithmic interpretation of this procedure explicit. Our framework offers a pedagogical and experimental testbed to explore how transformers use informational geometry to implement next-token prediction.
| Comments: | 27 pages, 15 figures, 2 tables. Code and training-dynamics animations: this https URL |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.09838 [cs.LG] |
| (or arXiv:2610.09838v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09838 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Toviah Moldwin [view email]
[v1]
Wed, 7 Oct 2026 11:01:54 UTC (2,541 KB)
来源:arXiv:cs.LG · arxiv.org