arXiv:cs.LG· Jiho Lee, Jeongeun Park, Heayoun Choi, Taekyung Kim, Eunwoo Kim·· 4 小时前AI 评分37
SOUL:稀疏特征策略遗忘缓解 VLA 模型的状态幻觉
Sparse Feature Policy Unlearning Mitigates State Hallucination in Vision-Language-Action Models
AI 导读
针对 VLA 模型的状态幻觉问题,研究者提出 SOUL(Sparse feature pOlicy UnLearning)方法,通过稀疏自编码器识别幻觉失败与成功行为对应的稀疏特征,分别作为遗忘与保留目标,选择性遗忘策略知识中与状态幻觉相关的部分。在多种 VLA 架构的仿真与真实环境实验中,该方法显著减少幻觉失败、提升任务成功率,且未明显损害原有操作能力。
正文
Abstract:Vision-Language-Action (VLA) models have shown strong generalization in robotic manipulation by leveraging rich representations from pretrained vision-language models. However, their deployment in real-world environments remains limited by recurring unreliable behaviors. In this work, we study state hallucination, a recurring failure pattern in which a VLA continues acting as if an unrealized robot-object state had been achieved. Our analyses find that state hallucination coincides with weakened attention to task-relevant visual regions, and a mechanistic interpretation via sparse autoencoders reveals that hallucination-associated sparse features are activated when these failures occur. Based on this analysis, we propose SOUL (Sparse feature pOlicy UnLearning), which selectively unlearns policy knowledge associated with state hallucination behaviors, where sparse features identified from hallucination failures and successful behaviors serve as explicit forgetting and retention targets, respectively. Experiments across VLA architectures in simulated and real-world environments show that our method substantially reduces hallucinated failures and improves task success without substantially compromising the existing manipulation capabilities. These results suggest that interpretable feature analysis provides a practical basis for selectively modifying undesirable knowledge in robot policies.
| Subjects: | Robotics (cs.RO); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.09496 [cs.RO] |
| (or arXiv:2610.09496v1 [cs.RO] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09496 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jiho Lee [view email]
[v1]
Wed, 7 Oct 2026 05:52:02 UTC (10,784 KB)
来源:arXiv:cs.LG · arxiv.org