arXiv:cs.LG(机器学习,全量分类)· Hyunsik Kim, Youngmoon Jung·· 15 小时前AI 评分45
Universal Byte-Level Encoding(UBE):用 UTF-8/UTF-16 双字母表路由缩小跨文字 token 预算差距
Universal Byte-Level Encoding: UTF-8/UTF-16 Routing to Reduce Cross-Script Token-Budget Disparities
AI 导读
研究者提出 Universal Byte-Level Encoding(UBE),一种双字母表 tokenizer:1-2 字节 UTF-8 字符走 UTF-8 路径,3-4 字节字符改走 UTF-16,从而降低 3 字节 BMP 字符的编码下限,且不抬高混合文字中英文片段的成本。
正文
Abstract:Byte-level byte-pair encoding (BBPE) tokenizers are attractive for multilingual large language models (LLMs) because they cover all Unicode text. In UTF-8-based BBPE, however, many scripts start from a higher fallback cost than English: when no learned merges can be applied, a multibyte character requires multiple byte-derived symbols. We call this worst-case pre-merge cost the encoding floor. A higher floor can increase token counts and per-request cost and shrink usable context. Changing the text encoding can reduce this gap, but a single global encoding can make already-efficient English spans more expensive in mixed-script text. We propose Universal Byte-Level Encoding (UBE), a dual-alphabet tokenizer that keeps 1-2-byte UTF-8 characters on the UTF-8 path while routing 3-4-byte UTF-8 characters through UTF-16. This lowers the encoding floor for 3-byte Basic Multilingual Plane (BMP) characters in scripts with high token premiums (token counts relative to English) without raising it for already-efficient spans in mixed-script text. UBE changes only the byte representation presented to byte-pair encoding (BPE); the merge rule remains standard, and exact decoding is preserved. UBE also composes with alternative boundary policies and morphology-based representations. In a Unicode 17 audit, UBE exactly round-trips all Unicode scalar values and all inputs in the official normalization, grapheme-break, and emoji test suites. Across intrinsic evaluations, UBE lowers dispersion in English-normalized token-count ratios, reducing cross-lingual token-budget disparity. In multilingual language model (LM) experiments, UBE matches BBPE's LM quality. In the main multilingual settings, UBE reduces token counts most for high-premium scripts and slightly lowers English token counts, yielding more usable context under fixed token budgets and faster prompt processing in content-matched benchmarks.
| Comments: | Accepted to NeurIPS 2026 |
| Subjects: | Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.01984 [cs.CL] |
| (or arXiv:2610.01984v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.01984 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Hyunsik Kim [view email]
[v1]
Thu, 1 Oct 2026 16:26:10 UTC (177 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org