Argo-Bench 发布:面向企业级工作流评测数据智能体
Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
论文提出 Argo-Bench,一个包含 210 项数据科学与分析任务的评测框架,基于纽约外卖平台模拟器构建 235 张表、75 亿行的 ERP 仓库(仿 Oracle E-Business Suite schema),要求智能体在仓库中导航重建事实并执行封禁欺诈账户、分配骑手激励预算等动作,由模拟器按后果评分。
Published on Oct 1
Authors:
,
,
,
Abstract
Real-world enterprise data science and analytics workflows require reasoning across dozens of tables, performing statistical analyses, and acting on the results. Established text-to-SQL benchmarks evaluate query generation alone, and audits have found their answer keys frequently wrong. Because real enterprise warehouses are too sensitive to release, these benchmarks are built on public datasets where a business event fits in a single table. We introduce Argo-Bench, an evaluation framework comprising 210 data science and analytics tasks. Drawing on public data, peer-reviewed industry literature, and regulatory filings, we simulate a food delivery platform in New York City at true scale, with 81 million orders in 2024, grounded economics, fraud patterns, and marketplace incentives. We export this world to an ERP warehouse of 235 tables and 7.5 billion rows, modeled on the Oracle E-Business Suite schema. The simulator's ground-truth state is withheld from the warehouse the agent sees, so tasks require reconstructing facts by navigating the warehouse before acting on them. Argo-Bench goes beyond text-to-SQL: the agent files actions such as banning fraudulent accounts, allocating courier incentive budgets, or issuing back pay, and the grader scores each by its consequences in the simulator. Every task has an executable reference solution that demonstrates solvability using only the warehouse. The strongest of 14 frontier and open-weight models scores 95 or higher on only 34.8% of tasks and averages 59.5 points. We hope Argo-Bench drives progress toward agents that understand, navigate, and act within real data environments.
View arXiv page View PDF Project page GitHub Add to collection
Models citing this paper 0
No model linking this paper
Cite arxiv.org/abs/2610.02122 in a model README.md to link it from this page.
Datasets citing this paper 2
textql/Argo-Bench
Viewer • Updated about 2 hours ago • 3.91B • 318 • 2
textql/Decision-Bench
Spaces citing this paper 0
No Space linking this paper
Cite arxiv.org/abs/2610.02122 in a Space README.md to link it from this page.
Collections including this paper 0
No Collection including this paper
Add this paper to a collection to link it from this page.
来源:HuggingFace Daily Papers(社区热门论文) · huggingface.co