跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Taebong Kim, Youngsik Hong, Minsik Kim, Sunyoung Choi, Jaewon Jang, Minseo Kim·· 18 小时前AI 评分43

异构序列混合器堆叠中位置无关、组合关键:拉丁方构造的 Aether-7B-5Attn

Placement Is Free, Composition Is Not: The Latin Square as a Provably-Balanced Construction for Heterogeneous Sequence-Mixer Stacks

AI 导读

研究者提出 Aether-7B-5Attn,一个 6.59B 参数(约 2.98B 激活)的 MoE 模型,其 49 层以 7×7 拉丁方排列七种序列混合机制,保证每种机制在每行每列各出现一次。

正文

View PDF HTML (experimental)

Abstract:Since GPT, most Transformers have repeated the same attention mechanism at every layer. Yet this design is largely a convention rather than a tested conclusion. When multiple sequence mixers are combined in one stack, improvements may arise from mechanism choice, placement, or both, making causal attribution difficult. We introduce Aether-7B-5Attn, a 6.59B-parameter mixture-of-experts model ($\approx$2.98B active) whose 49 layers contain seven sequence-mixing mechanisms arranged as a $7\times7$ Latin square. Because each mechanism appears exactly once in every row and column, the design guarantees balanced exposure across depth while eliminating placement confounds. To evaluate this principle, we build a parameter-matched proxy with four mechanisms arranged as a $4\times4$ Latin square over sixteen layers, matched to 700.9M parameters and trained with eight seeds per arm. The results reveal a clear dissociation. Rearranging a distributed heterogeneous stack into a balanced periodic cycle changes validation loss by only 0.16\%, indicating that exact placement has little effect. In contrast, clustering the same mechanisms into contiguous depth bands incurs a 0.59\% penalty, while replacing the heterogeneous stack with a homogeneous one incurs a 1.68\% penalty. These results indicate that performance depends primarily on heterogeneous composition distributed across depth rather than on any particular permutation. We confirm this finding at 2.16$\times$ larger scale (1.514B parameters), where the homogeneous-stack penalty increases to 2.63\% and removing the SSM-family mechanism produces a 3.20\% degradation. We further report per-mechanism cost profiles, English and Korean evaluations, and a causal-safety audit of all 49 layers. We release model weights, training recipes, training code, logs, and architecture source code.
Comments: 19pages, 5 figures
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2609.20269 [cs.LG]
  (or arXiv:2609.20269v3 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.20269

arXiv-issued DOI via DataCite

Submission history

From: Youngsik Hong [view email]
[v1] Wed, 29 Jul 2026 06:39:03 UTC (56 KB)
[v2] Mon, 28 Sep 2026 02:33:43 UTC (56 KB)
[v3] Thu, 1 Oct 2026 00:52:09 UTC (56 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org