跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Hyunsik Kim, Youngmoon Jung·· 15 小时前AI 评分45

Universal Byte-Level Encoding(UBE):用 UTF-8/UTF-16 双字母表路由缩小跨文字 token 预算差距

Universal Byte-Level Encoding: UTF-8/UTF-16 Routing to Reduce Cross-Script Token-Budget Disparities

AI 导读

研究者提出 Universal Byte-Level Encoding(UBE),一种双字母表 tokenizer:1-2 字节 UTF-8 字符走 UTF-8 路径,3-4 字节字符改走 UTF-16,从而降低 3 字节 BMP 字符的编码下限,且不抬高混合文字中英文片段的成本。

正文

View PDF HTML (experimental)

Abstract:Byte-level byte-pair encoding (BBPE) tokenizers are attractive for multilingual large language models (LLMs) because they cover all Unicode text. In UTF-8-based BBPE, however, many scripts start from a higher fallback cost than English: when no learned merges can be applied, a multibyte character requires multiple byte-derived symbols. We call this worst-case pre-merge cost the encoding floor. A higher floor can increase token counts and per-request cost and shrink usable context. Changing the text encoding can reduce this gap, but a single global encoding can make already-efficient English spans more expensive in mixed-script text. We propose Universal Byte-Level Encoding (UBE), a dual-alphabet tokenizer that keeps 1-2-byte UTF-8 characters on the UTF-8 path while routing 3-4-byte UTF-8 characters through UTF-16. This lowers the encoding floor for 3-byte Basic Multilingual Plane (BMP) characters in scripts with high token premiums (token counts relative to English) without raising it for already-efficient spans in mixed-script text. UBE changes only the byte representation presented to byte-pair encoding (BPE); the merge rule remains standard, and exact decoding is preserved. UBE also composes with alternative boundary policies and morphology-based representations. In a Unicode 17 audit, UBE exactly round-trips all Unicode scalar values and all inputs in the official normalization, grapheme-break, and emoji test suites. Across intrinsic evaluations, UBE lowers dispersion in English-normalized token-count ratios, reducing cross-lingual token-budget disparity. In multilingual language model (LM) experiments, UBE matches BBPE's LM quality. In the main multilingual settings, UBE reduces token counts most for high-premium scripts and slightly lowers English token counts, yielding more usable context under fixed token budgets and faster prompt processing in content-matched benchmarks.
Comments: Accepted to NeurIPS 2026
Subjects: Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as: arXiv:2610.01984 [cs.CL]
  (or arXiv:2610.01984v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.01984

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Hyunsik Kim [view email]
[v1] Thu, 1 Oct 2026 16:26:10 UTC (177 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org