arXiv:cs.CL· Mikey Watts (Independent Researcher), Yuchen Cui (University of California, Los Angeles)·· 3 小时前AI 评分52
Rephrase Before You Act:刻画并缓解视觉-语言-动作模型的指令措辞敏感性
Rephrase Before You Act: Characterizing and Mitigating Language Sensitivity in Vision-Language-Action Models
AI 导读
arXiv 论文(arXiv:2610.10526)指出视觉-语言-动作(VLA)模型对指令措辞高度敏感,一个词的改动可使 π0.5 在 LIBERO 开炉任务上的成功率从 100% 降到 2%。
正文
Authors:Mikey Watts (Independent Researcher), Yuchen Cui (University of California, Los Angeles)
Abstract:Vision-language-action models (VLAs) are strikingly sensitive to instruction phrasing and do not inherit the language robustness of the vision-language models they are built on. A one-word edit can move success by tens of points: $\pi_{0.5}$ turns on a LIBERO stove 100% of the time for "switch on the stove" and 2% for "switch on the hot plate", and a $\pi_0$ checkpoint finetuned with rephrase augmentation still shows swings of up to 61 points. We characterize this sensitivity with statistically tested single-edit swings and an oracle phrase search, which shows that phrasing alone nearly closes the 21-point gap between in-distribution and out-of-distribution tasks. We then reduce it without modifying the policy. Because the sensitivity is systematic, it can be expressed as explicit rules: we score many phrasings of a few training tasks, have a large language model distill the evidence into ten to twenty rephrasing rules, and at deployment rewrite each incoming instruction once under these rules. The rules improve the frozen $\pi_0$ by 16 to 27% relative on twelve held-out tasks across adversarial, VLM-generated, and human-generated phrasings, with gains concentrated on out-of-distribution tasks. The pipeline replicates on $\pi_{0.5}$ and LIBERO, lifting in-finetune success from 93.6% to 97.8%. The method requires no retraining and no per-step verification, and applies zero-shot to unseen tasks and instructions. Project website: this https URL
| Comments: | 9 pages, 8 figures, 3 tables. Project page: this https URL |
| Subjects: | Robotics (cs.RO); Computation and Language (cs.CL); Machine Learning (cs.LG) |
| ACM classes: | I.2.9; I.2.7 |
| Cite as: | arXiv:2610.10526 [cs.RO] |
| (or arXiv:2610.10526v1 [cs.RO] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10526 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Mikey Watts [view email]
[v1]
Wed, 7 Oct 2026 17:57:38 UTC (1,951 KB)
来源:arXiv:cs.CL · arxiv.org