跳到正文
arXiv:cs.AI· Ruihong Shen, \v{Z}iga Kova\v{c}i\v{c}, Peter Kulits, Xingrui Wang, Zizhang Li, Joshua B. Tenenbaum, Alan Yuille, Jieneng Chen, Jiajun Wu·· 4 小时前AI 评分39

4DCodeBench:在动态场景逆图形任务上评测 AI 智能体

4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes

AI 导读

4DCodeBench 是一个通过代码生成评测 4D 逆图形的基准,要求智能体从视频重建动态场景并输出可执行的图形程序。该基准汇集真实视频与合成场景,覆盖形变、流体、断裂等物理现象。对前沿模型的评测显示,较强的静态重建能力尚不能转化为对复杂动态的可靠重建。

正文

View PDF HTML (experimental)

Abstract:We introduce 4DCodeBench, a benchmark for 4D inverse graphics through code generation, in which agents reconstruct dynamic scenes from video as executable graphics programs. To accomplish this, agents must translate visual observations into compact representations of scene structure and dynamics, by implementing abstractions such as physical simulations to reproduce complex behavior. To evaluate this capability, we curate a set of real-world videos and construct synthetic scenes spanning diverse physical phenomena, including deformation, fluid flow, and fracture. We perform extensive benchmarking of frontier models, finding that strong static reconstruction capabilities do not yet translate into reliable reconstruction of complex dynamics. 4DCodeBench provides a testbed for tracking progress toward agents that can interpret the dynamics of the world through code. Our benchmark is available at this https URL
Comments: this https URL
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Graphics (cs.GR)
Cite as: arXiv:2610.03715 [cs.CV]
  (or arXiv:2610.03715v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2610.03715

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Peter Kulits [view email]
[v1] Fri, 2 Oct 2026 17:58:49 UTC (16,385 KB)

来源:arXiv:cs.AI · arxiv.org