跳到正文
arXiv:cs.AI· Lucas Bandarkar, Clark Peng, Ahmed Haj Ahmed, Aditi Khandelwal, Nanyun Peng·· 4 小时前AI 评分34

用 MoE 路由器实现纯解码器模型的跨语言对齐

Cross-Lingual Alignment for Decoder-Only Models using MoE Routers

AI 导读

研究者提出用 MoE 路由器的输出作为跨语言对齐目标,替代在隐藏状态上加辅助对齐损失的做法,因为路由器输出更适合对大量 token 做池化、实现序列级跨语言比较。在四个开源 MoE 上的受控持续预训练实验显示,加入该路由损失还能对齐底层隐藏表示,并在多样化评测集上提升多语言性能。

正文

View PDF HTML (experimental)

Abstract:Cross-lingual contrastive learning has been a core component of multilingual encoder training, but the ability to explicitly align representations is not possible in decoder-only LLMs because of varying multilingual tokenization. However, a growing amount of research suggests that even in LLMs, higher cross-lingual representational alignment leads to improved cross-lingual transfer. In this paper, we propose a novel approach to reimagine cross-lingual contrastive learning given the architectural constraints of modern LLMs. Rather than applying an auxiliary alignment loss on hidden states, we propose using the outputs of the mixture-of-experts (MoE) routers as the target for alignment. Router outputs lend themselves better to pooling over many tokens, enabling more reliable cross-lingual comparisons at the sequence-level. Controlled continual pre-training experiments on four open-source MoEs show that incorporating this routing loss also aligns the underlying hidden representations across languages. Most importantly, this loss improves multilingual performance on our diverse evaluation suite, demonstrating the potential of cross-lingual MoE router alignment.
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.01921 [cs.CL]
  (or arXiv:2610.01921v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.01921

arXiv-issued DOI via DataCite

Submission history

From: Lucas Bandarkar [view email]
[v1] Thu, 1 Oct 2026 15:56:54 UTC (4,320 KB)
[v2] Fri, 2 Oct 2026 12:35:42 UTC (4,320 KB)

来源:arXiv:cs.AI · arxiv.org