跳到正文
arXiv:cs.CL· Jiazheng Zhang, Long Ma, Yunxian Yang, Zhiheng Xi, Zhikai Lei, Yajie Yang, Chenyang Liao, Enyu Zhou, Yang Nan, Yuchen Tian, Senjie Jin, Yibo Wang, Wei He, Boyang Liu, Jixuan Huang, Xin Guo, Zhezheng Hao, Xinbing Liang, Zhihao Zhang, Changzhi Zhou, Wiggin Zhou, Tao Gui, Qi Zhang, Xuanjing Huang, Clarenceai, Aiden Adams·· 4 小时前

可验证环境扩展用于长周期工作智能体:WorkForge 论文被撤稿

Scaling Verifiable Environments for Long-horizon Work Agents

AI 导读

论文《Scaling Verifiable Environments for Long-horizon Work Agents》已被作者 Jiazheng Zhang 撤稿,原因是投稿时部分共同作者尚未完成审阅并批准稿件及作者名单。

正文

This paper has been withdrawn by Jiazheng Zhang

Authors:Jiazheng Zhang, Long Ma, Yunxian Yang, Zhiheng Xi, Zhikai Lei, Yajie Yang, Chenyang Liao, Enyu Zhou, Yang Nan, Yuchen Tian, Senjie Jin, Yibo Wang, Wei He, Boyang Liu, Jixuan Huang, Xin Guo, Zhezheng Hao, Xinbing Liang, Zhihao Zhang, Changzhi Zhou, Wiggin Zhou, Tao Gui, Qi Zhang, Xuanjing Huang, Clarenceai, Aiden Adams

No PDF available, click to view other formats

Abstract:Work agents operate over digital artifacts to execute professional knowledge-intensive work, requiring training environments that support long-horizon interaction and trustworthy verification. However, hand-crafted environments incur prohibitive engineering overhead that prevents environment scaling, whereas synthesis methods sacrifice workspace complexity, realism, or grounded verifiability. To bridge this gap, we introduce WorkForge, a scalable synthesis framework for constructing verifiable work-agent environments from real-world resources. Starting from expert workflows, WorkForge first identifies the resources, decisions, and deliverables required by each workflow. It then retrieves relevant real-world files and organizes them into a workspace. WorkForge inspects the workspace to extract concrete, checkable facts about its content. These factual anchors fix which task types the workspace can support and how their outcomes can be verified. Therefore, WorkForge derives each task's instructions, solution plan, and complementary programmatic and semantic verifiers directly from these factual anchors, keeping verification traceable to observable workspace evidence. Furthermore, we construct 16.7K verifiable environments across 40 professional domains, with workspaces collectively covering 60 file types. Post-training Qwen3.5-35B-A3B-Base improves GDPVal from 45.5 to 73.6 and APEX Score from 5.0 to 21.3, while enabling Qwen3.5-27B to achieve highly competitive performance and outperform strong competitors. Our analyses confirm the efficacy of the proposed method and reveal consistent scaling behaviors across both data volume and interaction horizons.
Comments: This version was submitted before all co-authors had completed their review and approved the manuscript and author list. We are withdrawing it while these issues are resolved
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.04906 [cs.CL]
  (or arXiv:2610.04906v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.04906

arXiv-issued DOI via DataCite

Submission history

From: Jiazheng Zhang [view email]
[v1] Sun, 4 Oct 2026 03:36:14 UTC (2,187 KB)
[v2] Thu, 8 Oct 2026 06:17:17 UTC (1 KB) (withdrawn)

来源:arXiv:cs.CL · arxiv.org