arXiv:cs.AI· Yoshinari Fujinuma, Keisuke Kamahori, Ryuto Koike, Abdelrahman Madkour, Varun Prashant Gangal, Monty Bichouna, Martyna Markiewicz, Shivani Jain, Duncan Curtis, Rebecca Qian, Anand Kannappan·· 10 小时前AI 评分48
SpeedrunBench:用电子游戏速通挑战 LLM 智能体
SpeedrunBench: Challenging LLM Agents with Video Game Speedrunning
AI 导读
研究者推出 SpeedrunBench,一个让前沿 LLM 智能体在 9 款电子游戏速通中比拼的基准,要求智能体反复改进策略、反思表现并长程推理,不断超越自己和他人。实验显示,前沿智能体在简单平台跳跃游戏中接近人类世界纪录,但在更长、更复杂的游戏和实际预算下仍落后于人类。该基准几乎总有更快的完成时间等待被发现,因此可作为抗饱和的智能体策略形成能力评测。
正文
Authors:Yoshinari Fujinuma, Keisuke Kamahori, Ryuto Koike, Abdelrahman Madkour, Varun Prashant Gangal, Monty Bichouna, Martyna Markiewicz, Shivani Jain, Duncan Curtis, Rebecca Qian, Anand Kannappan
Abstract:Frontier LLM agents have been shown to be capable of solving increasingly complex tasks for which humans have measurable solutions. This begs the pertinent question of whether LLM agents can go beyond what humans have already solved. The ability to develop sophisticated strategies to tackle consequential problems becomes paramount as well-trodden, human-developed solutions become insufficient for problems for which we lack context or enough training data. We study agents' capability of such strategy formation through the communal practice of video game speedrunning. In speedrunning, practitioners compete to find the fastest way to complete a video game under certain conditions, and in so doing uncovering interesting unorthodox play styles that require a thorough understanding and mastery of the underlying game mechanics. We introduce SPEEDRUNBENCH, a benchmark that evaluates frontier LLM agents across 9 different games. To perform well in this benchmark, agents must repeatedly improve their strategy, reflect on their performance, exploit their gained knowledge, and reason across a long-horizon of actions to improve on an increasingly difficult problem: being faster than themselves and everyone else. Our experiments show that while frontier agents approach human world records in simple platformer games, they remain behind human performance on longer, more complex games under practical budgets. These results suggest that SPEEDRUNBENCH is a useful testbed for studying agents' strategy formation capabilities as well as being a saturation-resistant evaluation measure, as there is almost always a faster completion time waiting to be discovered.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.08076 [cs.AI] |
| (or arXiv:2610.08076v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08076 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yoshinari Fujinuma [view email]
[v1]
Tue, 6 Oct 2026 10:06:25 UTC (4,265 KB)
来源:arXiv:cs.AI · arxiv.org