跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Animesh Varma·· 14 小时前AI 评分33

STATERA:用冻结时序 tubelet 零样本 sim-to-real 估计隐藏质量

STATERA: Hidden Mass Estimation via Zero-Shot Sim-to-Real Kinematics using Frozen Temporal Tubelets

AI 导读

STATERA 通过冻结大部分权重的 V-JEPA 视频骨干加轻量时序 tubelet mixer,从单目短视频预测不透明非对称刚体的逐帧质心热力图与轨迹。团队同时发布 HiddenMass Benchmark,含 50K 条 MuJoCo 轨迹和 63 段带物理标定质心真值的真实测试序列。

正文

View PDF HTML (experimental)

Abstract:Vision models pretrained for frame-level appearance often struggle to infer hidden physical properties from motion. We study center-of-mass (CoM) localization for opaque, asymmetric rigid bodies from short monocular videos, where surface cues and point tracking are unreliable under self-occlusion. We propose STATERA, which adapts a pretrained video backbone (V-JEPA) with mostly frozen weights and a lightweight temporal tubelet mixer to predict per-frame CoM heatmaps and trajectories. To support this task, we introduce the HiddenMass Benchmark, comprising 50K MuJoCo trajectories and a 63-sequence real-world test set with physically calibrated CoM ground truth. In simulation, STATERA-50K-Sigma improves normalized CoM error from 41.7% (DINOv2) to 25.2%. In zero-shot sim-to-real transfer, we observe a fundamental trade-off in supervision: phase-aware targets can induce bimodal predictions, while phase-agnostic targets can collapse toward statistically safe centroids. Nevertheless, our phase-aware STATERA-50K-Crescent is the only evaluated method that demonstrates consistent movement toward the true hidden offset. While this leads to a monocular vector overshoot artifact that marginally increases absolute Euclidean error compared to a static geometric centroid, it improves physics capture from 2.6% to 41.0%. These results suggest that frozen temporal representations can better separate inertial dynamics from visual geometry for hidden-parameter estimation.
Comments: 17 pages, 7 figures, 3 tables. Preprint
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Robotics (cs.RO)
Cite as: arXiv:2610.00003 [cs.CV]
  (or arXiv:2610.00003v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2610.00003

arXiv-issued DOI via DataCite

Submission history

From: Animesh Varma [view email]
[v1] Thu, 21 May 2026 20:25:20 UTC (2,795 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org