arXiv:cs.AI· Dheeraj Varghese, Anna Vettoruzzo, Walter Simoncini, Michelle Lorena Acevedo Callejas, Mohammad Mahdi Derakhshani, Kristof Meding, Joaquin Vanschoren, Cees G. M. Snoek·· 6 小时前AI 评分58
PlaySuite:面向交互式视觉智能的大规模基准
PlaySuite: A Large-Scale Benchmark for Interactive Visual Intelligence
AI 导读
研究者推出 PlaySuite,一个基于 5K 多款开源游戏的大规模交互式视觉智能评测基准,覆盖 Pygame、HTML5、Godot、Unity 等引擎,并配套统一闭环交互框架和 Video-LLM-as-a-judge 评分协议。
正文
Abstract:Recent advances in multimodal foundation models yield strong performance on static perception and reasoning benchmarks, yet such evaluations largely overlook a central aspect of intelligence: acting competently in dynamic environments over extended time horizons. We introduce PlaySuite, a large-scale benchmark for evaluating interactive visual intelligence across more than 5K open-source video games curated from PyWeek and this http URL. Spanning diverse genres and engines, including Pygame, HTML5, Godot, and Unity, these independent games are largely out-of-distribution for current models, reducing the likelihood that success can be achieved by retrieving memorized walkthroughs or web-scale training artifacts. To enable scalable evaluation across heterogeneous titles, we develop a unified closed-loop interaction framework optimized for HPC clusters alongside a Video-LLM-as-a-judge protocol that maps observable gameplay milestones to standardized progress levels. We evaluate fourteen recent open models spanning vision-language models, computer-use agents, and vision-language-action models. Our results yield strong evidence of a perception-action gap: despite strong reasoning capabilities, current models struggle to make sustained progress and exhibit recurring failures in spatial grounding, action execution, and self-correction. PlaySuite provides a reproducible and extensible testbed for measuring progress from visual perception to goal-directed interaction, and a foundation for developing models that can act, adapt, and generalize in dynamic visual environments.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.07127 [cs.CV] |
| (or arXiv:2610.07127v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07127 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Dheeraj Varghese [view email]
[v1]
Mon, 5 Oct 2026 17:52:40 UTC (17,373 KB)
来源:arXiv:cs.AI · arxiv.org