跳到正文
arXiv:cs.AI· Bochen Yang, Lianlei Shan·· 4 小时前

PearlVLA:在潜空间中渐进式精炼具身动作规划

PearlVLA: Progressive Embodied Action-Plan Refinement in Latent Space

AI 导读

PearlVLA 是一个在潜空间中渐进式精炼动作规划的 VLA 框架,它利用冻结的潜世界模型 LaWM 预测每个中间规划的潜视觉子目标,再由未来引导的规划精炼器据此更新规划,经 K 轮迭代后一次性输出动作块。

正文

View PDF HTML (experimental)

Abstract:Current Vision-Language-Action (VLA) models face a trade-off between efficient action generation and explicit deliberation. Directly decoding actions from vision-language backbone representations enables low-latency control, whereas textual reasoning, pixel-level subgoals, or world-model evaluation of decoded actions can improve planning but incur substantial latency and computational cost. We propose PearlVLA, a VLA framework that progressively refines a VLM-derived latent plan using feedback from the predicted consequence of each intermediate plan. PearlVLA uses a frozen latent world model (LaWM) pretrained on action-free video. At each refinement round, the current plan produces a continuous latent action code, and the LaWM predicts the corresponding latent visual subgoal. A future-guided plan refiner uses this subgoal to update the plan, so each revision reshapes the next LaWM query. After K rounds, the refined plan is passed once to the host policy's action interface to produce an action chunk. We further introduce Causal Refinement-Grouped Process-Reward RL to optimize latent refinement by comparing rewards from the longer-horizon imagined futures of plan edits made at the same refinement state. Experiments on the LIBERO and RoboCasa benchmarks show that PearlVLA performs competitively against strong existing methods.
Comments: 21 pages, 3 figures. Preprint
Subjects: Robotics (cs.RO); Artificial Intelligence (cs.AI)
Cite as: arXiv:2606.17924 [cs.RO]
  (or arXiv:2606.17924v2 [cs.RO] for this version)
  https://doi.org/10.48550/arXiv.2606.17924

arXiv-issued DOI via DataCite

Submission history

From: Bochen Yang [view email]
[v1] Tue, 16 Jun 2026 13:38:03 UTC (1,492 KB)
[v2] Thu, 8 Oct 2026 14:24:06 UTC (2,514 KB)

来源:arXiv:cs.AI · arxiv.org