跳到正文
arXiv:cs.AI· Shufan Li, Konstantinos Kallidromitis, Akash Gokul, Yusuke Kato, Kazuki Kozuka, Aditya Grover·· 3 小时前

MobileWorldBench:面向移动智能体的语义世界建模基准

MobileWorldBench: Towards Semantic World Modeling For Mobile Agents

AI 导读

研究者提出 MobileWorldBench,用于评估视觉语言模型(VLM)作为移动 GUI 智能体世界模型的能力,并发布含 1.4M 样本的大规模数据集 MobileWorld。他们进一步提出将 VLM 世界模型接入移动智能体规划框架的方法,以自然语言描述状态转移而非预测原始像素,实验显示语义世界模型能通过提升任务成功率直接帮助移动智能体。代码与数据集已公开。

正文

View PDF HTML (experimental)

Abstract:World models have shown great utility in improving the task performance of embodied agents. While prior work largely focuses on pixel-space world models, these approaches face practical limitations in GUI settings, where predicting complex visual elements in future states is often difficult. In this work, we explore an alternative formulation of world modeling for GUI agents, where state transitions are described in natural language rather than predicting raw pixels. First, we introduce MobileWorldBench, a benchmark that evaluates the ability of vision-language models (VLMs) to function as world models for mobile GUI agents. Second, we release MobileWorld, a large-scale dataset consisting of 1.4M samples, that significantly improves the world modeling capabilities of VLMs. Finally, we propose a novel framework that integrates VLM world models into the planning framework of mobile agents, demonstrating that semantic world models can directly benefit mobile agents by improving task success rates. The code and dataset is available at this https URL
Comments: 21 pages, 13 figures
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2512.14014 [cs.AI]
  (or arXiv:2512.14014v2 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2512.14014

arXiv-issued DOI via DataCite

Submission history

From: Shufan Li [view email]
[v1] Tue, 16 Dec 2025 02:16:42 UTC (3,146 KB)
[v2] Wed, 7 Oct 2026 23:06:49 UTC (3,215 KB)

来源:arXiv:cs.AI · arxiv.org