arXiv:cs.AI· Tuan Duong Trinh, Basim Azam, Mohammed Ishaq Ansari, Mohammed Yaqoob Ansari, Naveed Akhtar·· 4 小时前
推理链作为视觉-语言-动作策略的控制面板:DeepThinkVLA 反事实干预研究
Altered Thoughts, Altered Actions: Reasoning Chain as Control Surface for a Vision-Language-Action Policy
AI 导读
研究以确定性实体替换干预 DeepThinkVLA,检验推理链被改写时对机器人动作的影响。在 LIBERO-Goal 上,策略接收被污染的指令但换回干净指令生成的推理链,可恢复 47.8 pp 的成功率,但全部 10 项任务的成功率仍比干净运行低 38.0 pp,超出此前预测的 5-15 pp 差距。
正文
Abstract:Vision-language-action policies map camera images and natural-language instructions to a robot's motor actions. Some of these policies are designed to reason in text before acting, generating a reasoning chain and decoding actions conditioned on that chain. The works introducing this design offer the reasoning chain as an oversight interface: text a person can read and edit to correct the policy. What an edited reasoning chain does to the policy's motor actions, whether it repairs them or corrupts them, has so far been measured only in part. We measure both directions, repair and corruption, with our deterministic entity swap applied to the instruction the policy receives and to the reasoning chain it generates. A forty-task observed backdrop across all four LIBERO simulation suites reveals that the cost of corrupting the reasoning chain concentrates where language alone determines the goal. There, on LIBERO-Goal, we run the counterfactual intervention with DeepThinkVLA, chosen because its reasoning chain is exposed as plain text. The policy receives a corrupted instruction, but its reasoning chain is replaced by the one it generates when that instruction is clean. This counterfactually correct reasoning chain recovers 47.8 pp of the lost success, our pre-registered confirmatory test. Had the chain merely restated what the camera image already determines, the replacement could have changed nothing. Instead, all 10 tasks move in the predicted direction, though success falls short of the clean runs by 38.0 pp, a gap we had predicted at 5-15 pp. The reasoning chain is therefore a working control surface: text written into it moves the robot, repairing behaviour when the text is right and corrupting it when the text is wrong. Whether to expose such a control surface is a deployment tradeoff, and part of it can now be measured.
| Comments: | Accepted at the VLM4RWD workshop, NeurIPS 2026 |
| Subjects: | Robotics (cs.RO); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| MSC classes: | 68T40 |
| ACM classes: | I.2.9; I.2.6 |
| Cite as: | arXiv:2603.12717 [cs.RO] |
| (or arXiv:2603.12717v3 [cs.RO] for this version) | |
| https://doi.org/10.48550/arXiv.2603.12717 arXiv-issued DOI via DataCite |
Submission history
From: Tuan Duong Trinh [view email]
[v1]
Fri, 13 Mar 2026 07:02:51 UTC (168 KB)
[v2]
Tue, 29 Sep 2026 08:05:36 UTC (2,470 KB)
[v3]
Thu, 8 Oct 2026 06:15:00 UTC (2,474 KB)
来源:arXiv:cs.AI · arxiv.org