arXiv:cs.LG(机器学习,全量分类)· Animesh Varma·· 14 小时前AI 评分33
STATERA:用冻结时序 tubelet 零样本 sim-to-real 估计隐藏质量
STATERA: Hidden Mass Estimation via Zero-Shot Sim-to-Real Kinematics using Frozen Temporal Tubelets
AI 导读
STATERA 通过冻结大部分权重的 V-JEPA 视频骨干加轻量时序 tubelet mixer,从单目短视频预测不透明非对称刚体的逐帧质心热力图与轨迹。团队同时发布 HiddenMass Benchmark,含 50K 条 MuJoCo 轨迹和 63 段带物理标定质心真值的真实测试序列。
正文
Abstract:Vision models pretrained for frame-level appearance often struggle to infer hidden physical properties from motion. We study center-of-mass (CoM) localization for opaque, asymmetric rigid bodies from short monocular videos, where surface cues and point tracking are unreliable under self-occlusion. We propose STATERA, which adapts a pretrained video backbone (V-JEPA) with mostly frozen weights and a lightweight temporal tubelet mixer to predict per-frame CoM heatmaps and trajectories. To support this task, we introduce the HiddenMass Benchmark, comprising 50K MuJoCo trajectories and a 63-sequence real-world test set with physically calibrated CoM ground truth. In simulation, STATERA-50K-Sigma improves normalized CoM error from 41.7% (DINOv2) to 25.2%. In zero-shot sim-to-real transfer, we observe a fundamental trade-off in supervision: phase-aware targets can induce bimodal predictions, while phase-agnostic targets can collapse toward statistically safe centroids. Nevertheless, our phase-aware STATERA-50K-Crescent is the only evaluated method that demonstrates consistent movement toward the true hidden offset. While this leads to a monocular vector overshoot artifact that marginally increases absolute Euclidean error compared to a static geometric centroid, it improves physics capture from 2.6% to 41.0%. These results suggest that frozen temporal representations can better separate inertial dynamics from visual geometry for hidden-parameter estimation.
| Comments: | 17 pages, 7 figures, 3 tables. Preprint |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Robotics (cs.RO) |
| Cite as: | arXiv:2610.00003 [cs.CV] |
| (or arXiv:2610.00003v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00003 arXiv-issued DOI via DataCite |
Submission history
From: Animesh Varma [view email]
[v1]
Thu, 21 May 2026 20:25:20 UTC (2,795 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org