跳到正文
arXiv:cs.CL· Yang Liu, Jiaqi Li, Jun Bai, Qianyu Yang, Xiaobo Hu, Tao Peng, Zaiyuan Wang, Ran Tian, Jiayun Dong, Chun Zhang, Zixia Jia, Kaiyuan Chen, Yixin Ren, Yang Liu, Yanglihong Xiao, Lingyue Yin, Tiliang Duan, Ge Zhang, Gang Yao, Hao Chen, Yuan Gong, Jianpeng Jiao, Zilong Zheng·· 3 小时前

$OneMillion-Bench:语言智能体离人类专家还有多远?

\$OneMillion-Bench: How Far are Language Agents from Human Experts?

AI 导读

研究者提出 $OneMillion-Bench($OMB),一个覆盖法律、金融、工业、医疗和自然科学五大领域、共 400 项专家编写任务的基准,用于评测语言智能体在经济相关场景中的表现。

正文

Authors:Yang Liu, Jiaqi Li, Jun Bai, Qianyu Yang, Xiaobo Hu, Tao Peng, Zaiyuan Wang, Ran Tian, Jiayun Dong, Chun Zhang, Zixia Jia, Kaiyuan Chen, Yixin Ren, Yang Liu, Yanglihong Xiao, Lingyue Yin, Tiliang Duan, Ge Zhang, Gang Yao, Hao Chen, Yuan Gong, Jianpeng Jiao, Zilong Zheng

View PDF HTML (experimental)

Abstract:As language models (LMs) evolve from chat assistants to long-horizon agents capable of multi-step reasoning and tool use, existing benchmarks remain largely confined to structured or exam-style tasks that fall short of real-world professional demands. To this end, we introduce \$OneMillion-Bench (\$OMB), a benchmark of 400 expert-curated tasks spanning Law, Finance, Industry, Healthcare, and Natural Science, built to evaluate agents across economically consequential scenarios. Unlike prior work, the benchmark requires retrieving authoritative sources, resolving conflicting evidence, applying domain-specific rules, and making constraint decisions, where correctness depends as much on the reasoning process as the final answer. We adopt a rubric-based evaluation protocol scoring factual accuracy, logical coherence, practical feasibility, and professional compliance, focusing on expert-level problems to ensure meaningful differentiation across agents. Together, \$OMB provides a unified testbed for assessing agentic reliability, professional depth, and an indicator of practical readiness in domain-intensive scenarios.
Comments: NeurIPS 2026 (Evaluations and Datasets Track); the data and code is available at this https URL
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as: arXiv:2603.07980 [cs.LG]
  (or arXiv:2603.07980v2 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2603.07980

arXiv-issued DOI via DataCite

Submission history

From: Yang Liu [view email]
[v1] Mon, 9 Mar 2026 05:32:42 UTC (1,664 KB)
[v2] Thu, 8 Oct 2026 15:02:13 UTC (4,665 KB)

来源:arXiv:cs.CL · arxiv.org