arXiv:cs.CL· Clara Meister, G\"ul Sena Alt{\i}nta\c{s}, Antoine Bosselut·· 3 小时前
多语言语言模型中分词器选择的语言特异性影响
Language-Specific Effects of Tokenizer Choice in Multilingual Language Models
AI 导读
研究训练了 123 个语言模型、覆盖 54 种分词器,固定架构、训练语料、训练 token 预算与优化设置,仅改变分词器。结果显示,语言模型训练数据越少的语言,其 BPB 在各分词器间的标准差越大(31 种训练语言上 Spearman rho = -0.52),分词器训练中被排除的语言在所有研究语言上 BPB 均上升。
正文
Abstract:Tokenizer choice affects multilingual language modeling, but vocabulary capacity is finite and vocabulary size is often constrained: improving representation for some languages often comes at the expense of others. We therefore ask whether tokenizer choice matters equally across languages, a question that the current literature leave unanswered. To this end, we train 123 language models spanning 54 tokenizers. In the main comparison, architecture, training corpus, training-token budget, and optimization are held fixed, so the models differ only in their tokenizer. We find that tokenizer choice matters more for languages with less language-model training data: across the 54 tokenizers, the standard deviation of a language's bits-per-byte (BPB) increases as its model training-data share decreases (Spearman rho = -0.52 over the 31 trained languages and -0.69 over the 28 written with word boundaries). Leaving a language out of tokenizer training raises its BPB in every language we study, and the penalty tends to be larger for languages with less language-model training data. Giving lower-resource languages a larger share of tokenizer-training data, however, does not unconditionally help those languages: both equal weighting and an allocation inverting the shares with respect to the language model training data increase their BPB, particularly when language-model training repeats data. Finally, which intrinsic tokenizer properties are associated with better BPB differs across languages, providing further evidence that what makes a good tokenizer depends on the language. We find that the metrics quantifying these properties can be successfully used to predict downstream models' pairwise BPB rankings, suggesting a practical strategy for screening tokenizer candidates before training language models.
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.12144 [cs.CL] |
| (or arXiv:2610.12144v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.12144 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Clara Meister [view email]
[v1]
Thu, 8 Oct 2026 15:29:37 UTC (285 KB)
来源:arXiv:cs.CL · arxiv.org