跳到正文
arXiv:cs.CL· Seunghan Kim, Minyeong Choe, Hyunil Kim, Haehyun Cho·· 3 小时前AI 评分56

研究提出 Bridge Routing Heads 解释 LLM 多语言多跳推理机制

Bridge Routing Heads: Where Multilingual Multi-hop Reasoning Lives in LLMs

AI 导读

一篇 EMNLP 2026 论文在 Llama 3.1 70B 和 Qwen 2.5 72B 中识别出 Bridge Routing Heads(BRH),五个语言的头集合几乎互斥,Llama 3.1 70B 的平均 Jaccard 相似度仅 0.017,Qwen 2.5 72B 为 0.057。

正文

View PDF HTML (experimental)

Abstract:Multilingual LLMs answer the same multi-hop reasoning question across languages, but we lack a mechanistic account of whether they share an internal circuit. We identify Bridge Routing Heads (BRH) in two large multilingual LLMs through a three-stage pipeline. The resulting language-specific head sets exhibit near-complete mutual exclusivity across the five languages, with a mean Jaccard similarity of only 0.017 for Llama 3.1 70B and 0.057 for Qwen 2.5 72B, revealing language-idiosyncratic circuits. Ablating general BRH increases two-hop Negative Log-Likelihood (NLL) by 39-89x the random-head baseline, providing direct causal evidence of their role. Amplifying these heads in a failing target-language pass rescues up to 51.7% of cross-lingual failures, with no training. The two models share this dual-circuit pattern but allocate heads differently: Llama concentrates chaining in a large general pool, while Qwen leans on larger language-specific pools. Together these results show that activation-level intervention alone can recover correct answers from cross-lingual reasoning failures.
Comments: Accepted at EMNLP 2026
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2610.09733 [cs.CL]
  (or arXiv:2610.09733v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.09733

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Minyeong Choe [view email]
[v1] Wed, 7 Oct 2026 09:28:25 UTC (18,293 KB)

来源:arXiv:cs.CL · arxiv.org