arXiv:cs.LG· Naaisha Agarwal, Yihan Wu, Xin Guan, Ayaan Garg, Yikuan Hu, Mohan Li, Vincenzo Collura, Wang-Zhou Dai, Yao-Xiang Ding, Emanuele Sansone·· 5 小时前AI 评分43
AI 能理解折纸的语言吗?OrigamiBench 评测视觉语言模型的物理推理能力
Can AI Understand the Language of Origami?
AI 导读
研究者提出 OrigamiBench,用基于物理折叠动作的高层语言评测 AI 对折纸合成机制的程序化理解。实验显示,现代视觉语言模型仅靠扩大模型规模并不能可靠提升对物理变换的推理能力,且模型难以将程序化信息与视觉观察对齐,表明视觉与语言表征的整合仍然薄弱。
正文
Abstract:Building AI systems that can plan, act, and create in the physical world requires more than pattern recognition. Such systems must reason about the generative mechanisms and constraints governing physical processes, using structured representations that connect observations, actions, and their effects. Yet, many existing benchmarks study these capabilities separately, focusing either on visual recognition or on abstract symbolic or programmatic reasoning. Origami provides a natural testbed that integrates these abilities: constructing shapes through folds requires visual perception, reasoning about geometric and physical constraints, and sequential planning, while remaining sufficiently structured for systematic evaluation. We introduce OrigamiBench, a benchmark for evaluating programmatic understanding of the mechanisms underlying origami synthesis through a high-level language of physically grounded fold actions. Experiments with modern vision-language models reveal that scaling model size alone does not reliably improve reasoning about physical transformations. Moreover, models struggle to ground programmatic information in visual observations, suggesting that visual and language representations remain weakly integrated.
| Comments: | This version: "Can AI Understand the Language of Origami?" - different paper from v1 with different authors - NeurIPS LP4FM (Outstanding Runner-Up Award) v1: OrigamiBench: An Interactive Environment to Synthesize Flat-Foldable Origamis ICML LM4Plan (Oral) |
| Subjects: | Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2603.13856 [cs.LG] |
| (or arXiv:2603.13856v3 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2603.13856 arXiv-issued DOI via DataCite |
Submission history
From: Emanuele Sansone [view email]
[v1]
Sat, 14 Mar 2026 09:33:29 UTC (3,139 KB)
[v2]
Tue, 17 Mar 2026 17:36:55 UTC (3,138 KB)
[v3]
Fri, 2 Oct 2026 07:07:32 UTC (3,615 KB)
来源:arXiv:cs.LG · arxiv.org