跳到正文
arXiv:cs.CL· Felermino D. M. A. Ali, Millicent Ochieng, Ogbemi Ekwejunor-Etchie, Ade Famoti, Jacki O'Neill, Debjit Paul·· 3 小时前

Latent Core Tokenizer:为多语言分词压缩,但更有意义

Latent Core Tokenizer: Compress, but Meaningfully

AI 导读

Latent Core Tokenizer(LCT)是一种语言无关的分词方法,将结构发现与词表构建分离,先借助最小描述长度、基于熵的边界信号和形态句法约束识别可复用语言单元,再构建共享词表。

正文

View PDF HTML (experimental)

Abstract:Tokenizers are commonly optimized for compression, but a compact vocabulary does not necessarily distribute its capacity evenly across languages. We introduce the Latent Core Tokenizer (LCT), a language-agnostic approach that separates structural discovery from vocabulary construction. LCT uses Minimum Description Length, entropy-based boundary signals, and morphotactic constraints to identify reusable linguistic units before constructing a shared vocabulary. Across 104 languages with a 200K-token vocabulary, LCT achieves lower fertility and higher MorphScore than BPE, Unigram, and parity-aware BPE, while maintaining comparable cross-lingual disparity in tokenization cost. Across four multilingual downstream benchmarks, LCT improves aggregate score by 1.48, 1.83, and 2.00 points over BPE, Unigram, and parity-aware BPE, respectively. Our findings show that compression alone does not predict representation quality and highlight the importance of morphology-driven structural discovery and how frequency is used to allocate the final vocabulary across languages.
Comments: Under review
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2610.12376 [cs.CL]
  (or arXiv:2610.12376v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.12376

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Felermino Ali [view email]
[v1] Thu, 8 Oct 2026 17:32:32 UTC (1,861 KB)

来源:arXiv:cs.CL · arxiv.org