arXiv:cs.CL· Shu-Kai Hsieh, Da-Chen Lian·· 6 小时前AI 评分45
Language Unalignability:多语言 LLM 评估为何存在跨文化不可对齐概念
Language Unalignability: Why Some Concepts Resist Cross-Cultural Benchmark Evaluation
AI 导读
一篇立场论文提出,当前多语言 LLM 评测依赖的 Translation-Isomorphism Assumption(TIA)对语用标记、敬语、历时分层词等概念在原理上就不成立,并用 usage-cloud 框架形式化定义了 α-unalignability。
正文
Abstract:Current evaluation of multilingual Large Language Models (LLMs) rests on an implicit Translation-Isomorphism Assumption (TIA): that semantic structures across languages are congruent and mutually mappable without loss of information. We argue that this assumption is not merely violated in practice, but ill-posed in principle for a typologically identifiable class of concepts, including pragmatic markers, honorifics, and diachronically stratified terms. We formalize this failure using a usage-cloud framework, representing concepts as point sets of contextualized embeddings. We define $\alpha$-unalignability as the impossibility of any mapping that simultaneously preserves lexical faithfulness (centroid correspondence) and structural faithfulness (local neighborhood topology). We provide three layers of evidence. Behaviorally, we show that FLORES-200 translation failures are predicted by language family and resource class but not by script, and that LOBSTER reasoning scores vary by family. Mechanistically, we report a Representation-Intervention Gap (RIG) in a nine-model case study on Yami: the models' activations encode a regularity along which Yami groups with other low-resource and Austronesian languages, yet interventions on language-specific neurons show no demonstrated advantage over random masks: the regularity is visible but not usable by this intervention. Finally, we operationalize these findings into a multidimensional diagnostic profile: Cycle-Consistency, Pragmatic-Load Disagreement, Manifold-Curvature Mismatch, and RIG. We argue that collapsing cultural competence into a single scalar incentivizes "probabilistic flattening," and that recognizing the unalignable class is a precondition for AI that respects, rather than erases, cultural divergence. This suggests that multilingual alignment is not a single well-defined objective, but a set of mutually incompatible projections.
| Comments: | Position paper. 32 pages (10 pages main text), 6 figures, 12 tables |
| Subjects: | Computation and Language (cs.CL) |
| ACM classes: | I.2.7 |
| Cite as: | arXiv:2610.08303 [cs.CL] |
| (or arXiv:2610.08303v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08303 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Da-Chen Lian [view email]
[v1]
Tue, 6 Oct 2026 13:10:29 UTC (291 KB)
来源:arXiv:cs.CL · arxiv.org