arXiv:cs.AI· Bochen Yang, Lianlei Shan·· 4 小时前
PearlVLA:在潜空间中渐进式精炼具身动作规划
PearlVLA: Progressive Embodied Action-Plan Refinement in Latent Space
AI 导读
PearlVLA 是一个在潜空间中渐进式精炼动作规划的 VLA 框架,它利用冻结的潜世界模型 LaWM 预测每个中间规划的潜视觉子目标,再由未来引导的规划精炼器据此更新规划,经 K 轮迭代后一次性输出动作块。
正文
Abstract:Current Vision-Language-Action (VLA) models face a trade-off between efficient action generation and explicit deliberation. Directly decoding actions from vision-language backbone representations enables low-latency control, whereas textual reasoning, pixel-level subgoals, or world-model evaluation of decoded actions can improve planning but incur substantial latency and computational cost. We propose PearlVLA, a VLA framework that progressively refines a VLM-derived latent plan using feedback from the predicted consequence of each intermediate plan. PearlVLA uses a frozen latent world model (LaWM) pretrained on action-free video. At each refinement round, the current plan produces a continuous latent action code, and the LaWM predicts the corresponding latent visual subgoal. A future-guided plan refiner uses this subgoal to update the plan, so each revision reshapes the next LaWM query. After K rounds, the refined plan is passed once to the host policy's action interface to produce an action chunk. We further introduce Causal Refinement-Grouped Process-Reward RL to optimize latent refinement by comparing rewards from the longer-horizon imagined futures of plan edits made at the same refinement state. Experiments on the LIBERO and RoboCasa benchmarks show that PearlVLA performs competitively against strong existing methods.
| Comments: | 21 pages, 3 figures. Preprint |
| Subjects: | Robotics (cs.RO); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2606.17924 [cs.RO] |
| (or arXiv:2606.17924v2 [cs.RO] for this version) | |
| https://doi.org/10.48550/arXiv.2606.17924 arXiv-issued DOI via DataCite |
Submission history
From: Bochen Yang [view email]
[v1]
Tue, 16 Jun 2026 13:38:03 UTC (1,492 KB)
[v2]
Thu, 8 Oct 2026 14:24:06 UTC (2,514 KB)
来源:arXiv:cs.AI · arxiv.org