arXiv:cs.LG(机器学习,全量分类)· Rikhat Akizhanov (MBZUAI), Yangsong Zhang (MBZUAI), Nikolai Kaliazin (MBZUAI), Peter Wolf (ETH Z\"urich), Yoshihiko Nakamura (MBZUAI), Pascal Fua (EPFL), Fabio Pizzati (MBZUAI), Ivan Laptev (MBZUAI)·· 13 小时前AI 评分36
PACT:从视频端到端学习人体姿态、接触与力
PACT: End-to-End Learning of Human Pose, Contacts, and Forces from Video
AI 导读
PACT 是一个从单目视频端到端联合估计人体姿态、接触与接触力的模型,通过为人体重建基础模型加入可学习的接触力 token 和时序 Transformer,并用物理监督约束运动与交互力的一致性。团队还构建了数据标注流水线,并发布真实攀岩基准 ForceWall,含攀岩视频与力传感器采集的接触力真值。实验显示其在接触与力估计上达到 SOTA,优于分阶段重建方法,并能泛化到训练分布之外的交互。
正文
Abstract:Human motion, environmental contacts, and interaction forces are governed by common physical laws, yet existing approaches typically separate visual pose reconstruction from contact and force estimation. This separation limits joint reasoning and can propagate errors between stages. We introduce PACT, an end-to-end model that jointly learns to estimate human pose, contacts and contact forces from monocular video. Our approach augments a human reconstruction foundation model with learnable contact-force tokens and a temporal transformer that integrates visual features with world-space motion. Joint prediction heads refine human poses and estimate contacts and forces, while physics-based supervision encourages consistency between the reconstructed motion and interaction forces. To address the scarcity of force annotations, we develop a data annotation pipeline that combines contact labeling with physics-based motion and force optimization, producing training supervision from synthetic and real-world videos. We also introduce a real-world climbing benchmark ForceWall with climbing videos and corresponding ground-truth contact forces obtained from the force sensors. Experiments demonstrate state-of-the-art contact and force estimation, outperforming staged reconstruction approaches and generalizing to interactions beyond the training distribution. These results support end-to-end joint learning as an effective approach to recovering human motion and physical interactions from video.
| Comments: | 31 pages, 12 figures. Project page: this https URL |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.00451 [cs.CV] |
| (or arXiv:2610.00451v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00451 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Rikhat Akizhanov [view email]
[v1]
Wed, 30 Sep 2026 17:59:57 UTC (7,922 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org