跳到正文
arXiv:cs.LG· Eray Turkel, Mengsha Sun, Kartik Ayyar, Sean Dunigan, Jack Lu, Vlad Shcherban, Hsiang-Shun Shih, Xin Wang, Tiantian Zhang·· 3 小时前AI 评分45

OpenGameEval:在状态化游戏引擎中评测智能体编程与探索能力

OpenGameEval: Benchmarking Agentic Programming and Exploration in a Stateful Game Engine

AI 导读

OpenGameEval 是一个在 Roblox Studio 内评测智能体游戏开发的基准与评估框架,通过可复现的状态化游戏引擎会话运行语言模型,并用可执行检查对编辑场景和模拟游玩分别打分。

正文

View PDF HTML (experimental)

Abstract:We present OpenGameEval, a benchmark and evaluation framework for agentic game development inside Roblox Studio. It runs language models as agents in reproducible, stateful game-engine sessions and scores each run with executable checks, both on the edited scene and in a simulated play session. Most agentic coding benchmarks require exploration but score only final task success. OpenGameEval separates observation tools from editing tools in its eight-tool action space, so exploration can be measured directly. We measure the pass rates and exploration behavior of 13 frontier models on 84 human-curated core tasks, with 16 attempts per task.
The tasks are hard for current models. The best model solves 51.7% of tasks on a single attempt and 39.4% five times out of five, and no tested model solves six of the tasks. Models at the frontier reach similar pass rates by solving different tasks: splitting tasks by the kind of work they require spreads the top five by 5.0pp on script-authoring tasks and 12.5pp on scene-change tasks.
Exploration behavior predicts whether a run succeeds. Holding task and model fixed, a run that inspects every object a reference solution touches before acting on it passes 13.4pp more often than a run that inspects none of them on scene-only tasks, and 9.8pp more often on script-only tasks.
We release the task suite, its place files, the per-task annotations, a plugin that runs the tasks inside Roblox Studio, and an updated leaderboard under the MIT license at this https URL.
Comments: A shorter version appears at the NeurIPS 2026 Workshop on Evaluation of Interactive Agents
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.02563 [cs.LG]
  (or arXiv:2610.02563v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.02563

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Eray Turkel [view email]
[v1] Thu, 1 Oct 2026 22:56:10 UTC (1,875 KB)

来源:arXiv:cs.LG · arxiv.org