arXiv:cs.LG(机器学习,全量分类)· Youngjun Jun, Kyumin Choi, Youngmin Kim, Seonghyun Jin, Sunwoo Park, Jangho Park, Jong Chul Ye·· 14 小时前AI 评分33
DriftOPD:面向一步式 VLA 策略的序列级反向 KL 蒸馏
DriftOPD: Sequence-Level Reverse-KL Distillation for One-Step VLA Policies
AI 导读
DriftOPD 是一个无需教师模型、无需 rollout 的框架,用于对连续 VLA 动作专家进行序列级 on-policy 蒸馏。该方法将序列级反向 KL 散度分解为 chunk 级反向 KL 项与未来势能项,分别用一步式 drifting 目标和从离线演示学习的 Q 函数 critic 优化,仅靠离线数据和一步动作生成即可实现序列级优化。
正文
Abstract:Vision-Language-Action (VLA) models increasingly rely on action experts that generate short action chunks under receding-horizon control. While chunk-level training is convenient across robot embodiments, it optimizes local action likelihood without explicitly accounting for long-horizon task success. Sequence-level reinforcement learning can address this limitation, but typically requires policy rollouts and closed-loop interaction, which are costly for real-robot manipulation. We introduce DriftOPD, a teacher-free, rollout-free framework for sequence-level on-policy distillation of continuous VLA action experts. We show that the sequence-level reverse Kullback-Leibler (KL) divergence decomposes into a chunk-level reverse-KL term and a future-potential term that captures the long-horizon effect of the current action. DriftOPD optimizes these two terms using a one-step drifting objective and a Q-function critic learned from offline demonstrations, respectively, enabling sequence-level optimization with only offline data and one-step action generation. Across multiple VLA architectures in simulation and real-world manipulation, DriftOPD generally outperforms existing one-step distillation baselines while achieving task success performance comparable to multi-step teacher policies. These results demonstrate that long-horizon behavior can be effectively distilled into one-step VLA action experts without online interaction or a separate teacher.
| Comments: | Preprint |
| Subjects: | Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.00317 [cs.RO] |
| (or arXiv:2610.00317v1 [cs.RO] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00317 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jong Chul Ye [view email]
[v1]
Tue, 29 Sep 2026 02:55:12 UTC (6,646 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org