跳到正文
arXiv:cs.LG· A. Said Gurbuz, Ahmed Nassar, Sunghwan Hong, Marc Pollefeys, Peter W. J. Staar·· 5 小时前AI 评分54

DeskForge:用桌面环境密集监督训练计算机使用 Agent

DeskForge: Dense Supervision from Desktop Environments for Computer-Use Agents

AI 导读

ETH 等机构研究者发布 DeskForge,一个可控桌面环境,通过组合真实应用并变化应用状态、窗口布局、外观和分辨率,融合截图、accessibility tree 和窗口几何信息生成密集元素标注,构建了含 1.2M 标注桌面观测、159.7M 元素实例的 DeskForge-1M 数据集。

正文

View PDF HTML (experimental)

Abstract:Computer-use agents need to reliably ground action targets in complex desktop scenes, where multiple applications, overlapping windows, and visually similar controls compete for attention. Existing training data rarely pair such scenes with dense annotations or vary them in a controlled way. We introduce DeskForge, a controllable desktop environment that composes and explores real applications to generate large-scale supervision for computer-use agents. It varies application states, content, window layout, appearance, and resolution, and fuses screenshots, accessibility trees, and window geometry into dense element annotations while recording the outcome of each executed action. Using this environment, we construct DeskForge-1M, a corpus of 1.2M annotated desktop observations containing 159.7M element instances. We fine-tune four vision-language models on 200K grounding examples drawn from DeskForge-1M. All four improve across held-out desktop conditions and on all five external GUI grounding benchmarks; for Qwen3.5-4B, accuracy increases by 11.51 percentage points on ScreenSpot-Pro and 10.11 points on OSWorld-G. The gains also translate to long-horizon task completion: under a fixed planner, the fine-tuned action models solve more WebArena-Infinity and OpenApps tasks, with Qwen3.5-4B increasing from 31 to 50 of 119 tasks and from 3 to 15 of 100 tasks, respectively. These results show that controllable composition of real desktop environments provides a scalable source of supervision for improving both GUI grounding and long-horizon computer use. The framework code, the dataset, and the fine-tuned model are available from the project page: this https URL
Comments: 37 pages, 15 figures, 12 tables. Project page: this https URL
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2610.02320 [cs.CV]
  (or arXiv:2610.02320v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2610.02320

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Abdurrahman Said Gürbüz [view email]
[v1] Thu, 1 Oct 2026 18:00:05 UTC (43,694 KB)

来源:arXiv:cs.LG · arxiv.org