arXiv:cs.LG(机器学习,全量分类)· Keisuke Shirai, Tomohiro Motoda, Hanbit Oh, Ryoichi Nakajo, Roman Mykhailyshyn, Ryo Hanai, Shotaro Miwa, Yukiyasu Domae·· 9 小时前AI 评分40
TOAST:面向自回归视觉-语言-动作模型的随机机器人动作分词方法
TOAST: Stochastic Robot Action Tokenization for Autoregressive Vision-Language-Action Models
AI 导读
研究者提出 TOAST,一种随机动作分词方法,在策略训练中对同一量化动作序列采样多种分词方式,在不增加演示数据的前提下丰富离散监督信号。在 LIBERO 基准上,TOAST 稳定优于确定性分词,仅用 1/16 训练数据时成功率提升 6.8 个百分点;在四项真实机器人操作任务上平均成功率提升 15.8 个百分点。
正文
Abstract:Autoregressive Vision-Language-Action models often represent continuous robot actions as discrete token sequences, enabling action prediction with standard next-token objectives. FAST has substantially improved this representation by compactly encoding action containing diverse temporal frequencies into relatively few tokens. However, while such compression reduces the number of action tokens required for autoregressive prediction, it does not necessarily improve the efficiency of policy learning from limited demonstrations. In particular, FAST typically assigns a single deterministic tokenization to each quantized action sequence, although multiple token sequences can represent and decode to the same robot motion. We investigate whether exploiting this representational redundancy can improve policy learning. In this paper, we propose TOkenization of Action sequences with STochastic sampling (TOAST), a stochastic action tokenization method that samples alternative tokenizations of the same quantized action sequence during policy training. This diversifies the discrete supervision while preserving the underlying robot action and requires no additional demonstrations. Experiments on LIBERO show that TOAST consistently improves over its deterministic counterpart, with the improvement increasing as training data decreases, achieving a 6.8 point gain in success rate when only 1/16 of training data is available. Across four real-robot manipulation tasks, TOAST further improves mean success rate by 15.8 points over the deterministic counterpart. These results demonstrate the effectiveness of stochastic action tokenization for autoregressive robot policy learning, particularly when training data are limited.
| Comments: | Project page: this https URL |
| Subjects: | Robotics (cs.RO); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.00899 [cs.RO] |
| (or arXiv:2610.00899v1 [cs.RO] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00899 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Keisuke Shirai [view email]
[v1]
Thu, 1 Oct 2026 01:28:33 UTC (4,940 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org