跳到正文
arXiv:cs.AI· Damir Shodiev, Nikita Kachaev, Alexey K. Kovalev, Aleksandr I. Panov, Aleksei Staroverov·· 4 小时前

VLA Grounder:面向黑盒 VLA 模型的语言条件空间优化

VLA Grounder: Language-Conditioning Space Optimization for Black-Box VLA Models

AI 导读

研究者提出 VLA Grounder,通过优化语言条件空间而非更新动作权重来改进冻结的 VLA 策略:该方法将人类指令转化为简短的 VLA 接地命令,并结合物体外观、空间关系与目标接地线索,用稀疏任务完成奖励进行强化学习优化,下游 VLA 全程保持冻结。

正文

View PDF HTML (experimental)

Abstract:Vision-Language-Action (VLA) models are commonly treated as end-to-end action policies conditioned on natural-language task descriptions. In practice, however, their behavior often depends sharply on how the instruction is phrased, suggesting that language is not merely a task label but an optimizable conditioning input. We study whether frozen VLA policies can be improved by optimizing language space rather than updating action weights. Our method introduces a language-conditioning space policy that translates a human instruction into a short VLA-grounded command using object appearance, spatial relations, and target-grounding cues. The language-conditioning space policy is optimized with reinforcement learning from sparse task-completion rewards, while the downstream VLA remains fully frozen. Experiments on RL4VLA and VL-Think show that language-conditioning space optimization improves success on instruction-sensitive, symbolic, and multi-object manipulation tasks, demonstrating that language can serve as an optimizable variable for robot foundation models.
Website: this https URL
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2607.04517 [cs.AI]
  (or arXiv:2607.04517v2 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2607.04517

arXiv-issued DOI via DataCite

Submission history

From: Aleksei Staroverov [view email]
[v1] Sun, 5 Jul 2026 21:41:25 UTC (4,408 KB)
[v2] Wed, 7 Oct 2026 19:59:37 UTC (8,278 KB)

来源:arXiv:cs.AI · arxiv.org