跳到正文
arXiv:cs.LG· Nanhe Chen, Runqiu Yang, Jiawei Tang, Sichao Liu, Yuquan Wang·· 6 小时前AI 评分46

CoRP:紧凑机器人策略需要细粒度视觉表征

Compact Robot Policies Need Fine-Grained Visual Representations

AI 导读

研究团队构建了 CoRP(Compressed Representation Policy),一个仅含 48.9M 参数、不含视觉语言模型和视频生成先验的紧凑策略,在 LIBERO 上达到 97.0%,在 RoboTwin 2.0 Clean/Randomized 上达到 75.78%/73.36%,性能匹敌比其大 40.9-163.6 倍的系统。

正文

View PDF HTML (experimental)

Abstract:Multi-task manipulation policies differ in architecture, scale, and pretrained priors all at once, so published comparisons cannot attribute performance to any single component. We argue that most of it comes from the visual representation, and that parameter scale and generative priors are largely incidental. To test this, we build CoRP (Compressed Representation Policy), a deliberately compact policy (48.9M parameters, no vision-language model and no video-generative prior) that factorizes into a representation extractor and a flow-matching action generator. It reaches 97.0% on LIBERO and 75.78%/73.36% on RoboTwin 2.0 Clean/Randomized, matching systems 40.9-163.6x larger. Holding the action generator fixed, we then vary one extractor property at a time. Pretrained initialization is decisive: a random ViT-S/14 drops to 78.1% and an ImageNet ResNet-34 to 74.5% on LIBERO. Pretraining alone is not enough, as freezing the encoder costs 19.8 points. Compression matters as much: resampling each view to 48 tokens beats passing all patch tokens (97.0% vs 83.2%), and a variational information bottleneck over those tokens is worse than a hard token budget, cutting LIBERO-Goal from 95.8% to 33.0% by suppressing the instruction-dependent token selection the policy relies on. Language conditioning contributes only where the observation leaves the goal ambiguous (LIBERO-Goal: 9.2% to 95.8%), while on RoboTwin 2.0, where observations are unambiguous, removing it slightly improves success. Therefore, we argue that a compact policy works when its representation is pretrained, task-adapted, and compressed. Project page: this https URL
Comments: 35 pages, 21 figures, 8 tables
Subjects: Robotics (cs.RO); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2610.08183 [cs.RO]
  (or arXiv:2610.08183v1 [cs.RO] for this version)
  https://doi.org/10.48550/arXiv.2610.08183

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Sichao Liu [view email]
[v1] Tue, 6 Oct 2026 11:36:19 UTC (17,509 KB)

来源:arXiv:cs.LG · arxiv.org