arXiv:cs.LG· Eray Turkel, Mengsha Sun, Kartik Ayyar, Sean Dunigan, Jack Lu, Vlad Shcherban, Hsiang-Shun Shih, Xin Wang, Tiantian Zhang·· 3 小时前AI 评分45
OpenGameEval:在状态化游戏引擎中评测智能体编程与探索能力
OpenGameEval: Benchmarking Agentic Programming and Exploration in a Stateful Game Engine
AI 导读
OpenGameEval 是一个在 Roblox Studio 内评测智能体游戏开发的基准与评估框架,通过可复现的状态化游戏引擎会话运行语言模型,并用可执行检查对编辑场景和模拟游玩分别打分。
正文
Abstract:We present OpenGameEval, a benchmark and evaluation framework for agentic game development inside Roblox Studio. It runs language models as agents in reproducible, stateful game-engine sessions and scores each run with executable checks, both on the edited scene and in a simulated play session. Most agentic coding benchmarks require exploration but score only final task success. OpenGameEval separates observation tools from editing tools in its eight-tool action space, so exploration can be measured directly. We measure the pass rates and exploration behavior of 13 frontier models on 84 human-curated core tasks, with 16 attempts per task.
The tasks are hard for current models. The best model solves 51.7% of tasks on a single attempt and 39.4% five times out of five, and no tested model solves six of the tasks. Models at the frontier reach similar pass rates by solving different tasks: splitting tasks by the kind of work they require spreads the top five by 5.0pp on script-authoring tasks and 12.5pp on scene-change tasks.
Exploration behavior predicts whether a run succeeds. Holding task and model fixed, a run that inspects every object a reference solution touches before acting on it passes 13.4pp more often than a run that inspects none of them on scene-only tasks, and 9.8pp more often on script-only tasks.
We release the task suite, its place files, the per-task annotations, a plugin that runs the tasks inside Roblox Studio, and an updated leaderboard under the MIT license at this https URL.
| Comments: | A shorter version appears at the NeurIPS 2026 Workshop on Evaluation of Interactive Agents |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.02563 [cs.LG] |
| (or arXiv:2610.02563v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02563 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Eray Turkel [view email]
[v1]
Thu, 1 Oct 2026 22:56:10 UTC (1,875 KB)
来源:arXiv:cs.LG · arxiv.org