arXiv:cs.AI· Max Ku, Nok-Kan Law, Yu-Chien Tang, Shih-Ying Yeh, Ping Nie, Andy Zheng, Tat Hei Lai, Fei-Yueh Chen, Nikko Yu, Wei-Chieh Sun, Suzy Huang, Chiao-Wei Hsu, Chih-Chuan Huang, Chak-Wing Mak, Ho Yin Sam Ng, Edisy Kin Wai Chan, Min-Hung Chen, Ho Kei Cheng·· 5 小时前AI 评分45
世界编辑:在可执行世界中以递增深度进行干预
World Editing: Intervening on Executable Worlds at Increasing Depth
AI 导读
研究提出"世界编辑"概念,即在保留原有属性的前提下对已有可执行世界进行干预,并引入"干预深度"描述编辑对世界实体、动态与系统的耦合强度。团队基于 Minecraft 和 Terraria 的工业级游戏模组构建了 IGMWorld 与 IGMBench,包含 110 项任务和超 1.1K 条可执行状态与行为标准。
正文
Authors:Max Ku, Nok-Kan Law, Yu-Chien Tang, Shih-Ying Yeh, Ping Nie, Andy Zheng, Tat Hei Lai, Fei-Yueh Chen, Nikko Yu, Wei-Chieh Sun, Suzy Huang, Chiao-Wei Hsu, Chih-Chuan Huang, Chak-Wing Mak, Ho Yin Sam Ng, Edisy Kin Wai Chan, Min-Hung Chen, Ho Kei Cheng
Abstract:Interactive world models are increasingly capable of generating environments and acting within them, yet deliberately editing an existing executable world remains underexplored. We formulate world editing as intervening on an existing world while preserving properties that should remain unchanged, and introduce intervention depth as an axis describing how strongly an edit couples world entities, dynamics, and systems. We instantiate this capability through industry-grade game modding and introduce IGMWorld, together with IGMBench, a benchmark of 110 tasks and over 1.1K executable state and behavioral criteria across Minecraft and Terraria. The tasks span property, entity, dynamics, and system interventions and are evaluated through deterministic executability, behavioral, preservation, and visual checks. Frontier coding agents already exhibit substantial world-editing capability: the strongest configuration solves 78.2% of tasks under a strict task-level criterion, while criterion-level performance reaches 94.8%. Reliability generally decreases with intervention depth, and this pattern persists even among tasks with similar numbers of evaluation criteria. Most failed edits still build and load successfully, suggesting that the main difficulty is making the edited world behave as requested. Visual consistency remains a separate weakness, with all evaluated configurations below 50% joint visual pass rate. These results show that world editing is a distinct capability from world generation and interaction, and that executable games provide a practical testbed for studying it.
| Comments: | Preprint. Project page: this https URL |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.02331 [cs.AI] |
| (or arXiv:2610.02331v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02331 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Max W.F. Ku [view email]
[v1]
Thu, 1 Oct 2026 18:05:47 UTC (1,876 KB)
来源:arXiv:cs.AI · arxiv.org