跳到正文
arXiv:cs.AI· Zibin Dong, Yicheng Liu, Shiduo Zhang, Baijun Ye, Yifu Yuan, Fei Ni, Jingjing Gong, Xipeng Qiu, Hang Zhao, Yinchuan Li, Jianye Hao·· 3 小时前

ActionCodec:什么造就了好的动作 tokenizer

ActionCodec: What Makes for Good Action Tokenizers

AI 导读

论文提出 ActionCodec 动作 tokenizer,从 VLA 优化视角给出设计原则:最大化时间 token 重叠、最小化词表冗余、增强多模态互信息与 token 独立性。在 LIBERO 上,经 ActionCodec 微调的 SmolVLM2-2.2B 无需机器人预训练即达 95.5% 成功率,架构增强后升至 97.4%,创下无机器人预训练 VLA 模型的新 SOTA。

正文

Authors:Zibin Dong, Yicheng Liu, Shiduo Zhang, Baijun Ye, Yifu Yuan, Fei Ni, Jingjing Gong, Xipeng Qiu, Hang Zhao, Yinchuan Li, Jianye Hao

View PDF HTML (experimental)

Abstract:Vision-Language-Action (VLA) models leveraging the native autoregressive paradigm of Vision-Language Models (VLMs) have demonstrated superior instruction-following and training efficiency. Central to this paradigm is action tokenization, yet its design has primarily focused on reconstruction fidelity, failing to address its direct impact on VLA optimization. Consequently, the fundamental question of \textit{what makes for good action tokenizers} remains unanswered. In this paper, we bridge this gap by establishing design principles specifically from the perspective of VLA optimization. We identify a set of best practices based on information-theoretic insights, including maximized temporal token overlap, minimized vocabulary redundancy, enhanced multimodal mutual information, and token independence. Guided by these principles, we introduce \textbf{ActionCodec}, a high-performance action tokenizer that significantly enhances both training efficiency and VLA performance across diverse simulation and real-world benchmarks. Notably, on LIBERO, a SmolVLM2-2.2B fine-tuned with ActionCodec achieves a 95.5\% success rate without any robotics pre-training. With advanced architectural enhancements, this reaches 97.4\%, representing a new SOTA for VLA models without robotics pre-training. We believe our established design principles, alongside the released model, will provide a clear roadmap for the community to develop more effective action tokenizers.
Subjects: Robotics (cs.RO); Artificial Intelligence (cs.AI)
Cite as: arXiv:2602.15397 [cs.RO]
  (or arXiv:2602.15397v2 [cs.RO] for this version)
  https://doi.org/10.48550/arXiv.2602.15397

arXiv-issued DOI via DataCite

Submission history

From: Zibin Dong [view email]
[v1] Tue, 17 Feb 2026 07:07:15 UTC (8,933 KB)
[v2] Thu, 8 Oct 2026 15:35:42 UTC (3,423 KB)

来源:arXiv:cs.AI · arxiv.org