跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Youngjun Jun, Kyumin Choi, Youngmin Kim, Seonghyun Jin, Sunwoo Park, Jangho Park, Jong Chul Ye·· 14 小时前AI 评分33

DriftOPD:面向一步式 VLA 策略的序列级反向 KL 蒸馏

DriftOPD: Sequence-Level Reverse-KL Distillation for One-Step VLA Policies

AI 导读

DriftOPD 是一个无需教师模型、无需 rollout 的框架,用于对连续 VLA 动作专家进行序列级 on-policy 蒸馏。该方法将序列级反向 KL 散度分解为 chunk 级反向 KL 项与未来势能项,分别用一步式 drifting 目标和从离线演示学习的 Q 函数 critic 优化,仅靠离线数据和一步动作生成即可实现序列级优化。

正文

View PDF HTML (experimental)

Abstract:Vision-Language-Action (VLA) models increasingly rely on action experts that generate short action chunks under receding-horizon control. While chunk-level training is convenient across robot embodiments, it optimizes local action likelihood without explicitly accounting for long-horizon task success. Sequence-level reinforcement learning can address this limitation, but typically requires policy rollouts and closed-loop interaction, which are costly for real-robot manipulation. We introduce DriftOPD, a teacher-free, rollout-free framework for sequence-level on-policy distillation of continuous VLA action experts. We show that the sequence-level reverse Kullback-Leibler (KL) divergence decomposes into a chunk-level reverse-KL term and a future-potential term that captures the long-horizon effect of the current action. DriftOPD optimizes these two terms using a one-step drifting objective and a Q-function critic learned from offline demonstrations, respectively, enabling sequence-level optimization with only offline data and one-step action generation. Across multiple VLA architectures in simulation and real-world manipulation, DriftOPD generally outperforms existing one-step distillation baselines while achieving task success performance comparable to multi-step teacher policies. These results demonstrate that long-horizon behavior can be effectively distilled into one-step VLA action experts without online interaction or a separate teacher.
Comments: Preprint
Subjects: Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
Cite as: arXiv:2610.00317 [cs.RO]
  (or arXiv:2610.00317v1 [cs.RO] for this version)
  https://doi.org/10.48550/arXiv.2610.00317

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Jong Chul Ye [view email]
[v1] Tue, 29 Sep 2026 02:55:12 UTC (6,646 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org