arXiv:cs.LG· Zirong Song, Zheng Lu, Haoran Liao, Wanqi Zhong, Yunhe Ni, Lijie Wang, Xingjie Fan, Zhisheng Chen, Yantang Qu, Meijia Chen, Tianyu Xin, Yiming Li, Xiuying Chen·· 9 小时前AI 评分36
LayerRoute:面向视觉-语言-动作策略的动作条件化混合层路由
LayerRoute: Action-Conditioned Mixture-of-Layers Routing for Vision-Language-Action Policies
AI 导读
LayerRoute 是一种动作条件化的表示路由接口,可让 VLA 策略自适应访问 VLM 各层表示并复用早期动作表示。它在仿真与真实基准上持续提升 StarVLA-π 和 π0.5,LIBERO Long 最高提升 7.2,仅增加 0.31% / 3.87% 参数。消融实验验证了动作条件化层路由的有效性,路由分析显示动作层与任务设置间存在结构化分配模式。
正文
Authors:Zirong Song, Zheng Lu, Haoran Liao, Wanqi Zhong, Yunhe Ni, Lijie Wang, Xingjie Fan, Zhisheng Chen, Yantang Qu, Meijia Chen, Tianyu Xin, Yiming Li, Xiuying Chen
Abstract:Vision-Language-Action (VLA) policies leverage pretrained vision-language models (VLMs) to guide action generation for robot control. VLMs provide hierarchical visual-semantic representations that evolve across layers, from local visual geometry to abstract, language-aligned semantics; different manipulation tasks may therefore require different mixtures of layer representations. Meanwhile, the action module maintains intermediate representations that evolve throughout action computation and may provide useful information for subsequent decisions. However, existing VLA interfaces offer limited flexibility in representation access: VLM information is exposed through fixed layer assignments for each action layer, while intermediate action states are only propagated implicitly through residual streams without explicit reuse. We introduce LayerRoute, an action-conditioned representation routing interface that enables adaptive access to VLM layers and action representations. The Layer Mixture Router dynamically forms mixtures of cached VLM representations, while Action-State Reread reuses earlier action representations. Across diverse simulation and real-world benchmarks, LayerRoute consistently improves StarVLA-$\pi$ and $\pi_{0.5}$, achieving up to 7.2 gains on LIBERO Long with only 0.31% / 3.87% additional parameters. Ablation studies validate the benefit of action-conditioned layer routing, while routing analyses reveal structured allocation patterns across action layers and task settings.
| Comments: | 15 pages, 7 figures, 16 tables, including appendices |
| Subjects: | Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2609.06079 [cs.AI] |
| (or arXiv:2609.06079v3 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2609.06079 arXiv-issued DOI via DataCite |
Submission history
From: Zheng Lu [view email]
[v1]
Sat, 5 Sep 2026 13:13:07 UTC (570 KB)
[v2]
Thu, 17 Sep 2026 07:36:49 UTC (948 KB)
[v3]
Fri, 2 Oct 2026 10:05:01 UTC (948 KB)
来源:arXiv:cs.LG · arxiv.org