跳到正文
arXiv:cs.CL· Abdu Sallouh, Nicholas Popovi\v{c}, Michael F\"{a}rber·· 3 小时前

MSA 是否在 LLM 中主导阿拉伯语方言?一项表征层分析

Does Modern Standard Arabic (MSA) Dominate Arabic Dialects in LLMs? A Representation-Level Analysis

AI 导读

研究将 Shani 和 Basirat(2025)的语言主导性框架扩展到 26 种阿拉伯语变体,未发现 MSA 在 LLM 内部表征中构成主导性表示。方言表征形成密集且高度重叠的空间,归一化互信息相较差异更大的语言明显下降;最强可分性效应也不限于中间层,会随架构向更靠后的层偏移。

正文

View PDF HTML (experimental)

Abstract:Large language models (LLMs) often default to Modern Standard Arabic (MSA) when generating Arabic, even when prompted with dialectal Arabic. A natural explanation is that their internal representations are dominated by MSA. We test this hypothesis by adapting the language-dominance framework of Shani and Basirat (2025) (this https URL) to 26 Arabic varieties. Across layers and model families, we find no evidence that MSA acts as a dominant internal representation for Arabic dialects. Instead, dialect representations form a dense and highly overlapping space: normalized mutual information drops sharply compared to patterns reported for more distinct languages. Moreover, the strongest separability effects are not confined to intermediate layers, but can shift toward later layers depending on the architecture. These findings challenge a common interpretation of MSA-biased generation: output preference does not necessarily reveal internal representational dominance. Analyses of multilingual and dialectal LLMs should therefore distinguish generation bias from the geometry of internal representations.
Comments: EMNLP 2026
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2610.11510 [cs.CL]
  (or arXiv:2610.11510v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.11510

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Nicholas Popovič [view email]
[v1] Thu, 8 Oct 2026 08:46:47 UTC (6,947 KB)

来源:arXiv:cs.CL · arxiv.org