跳到正文
原文
HuggingFace Daily Papers(社区热门论文)·· 12 小时前AI 评分55

Argo-Bench 发布:面向企业级工作流评测数据智能体

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

AI 导读

论文提出 Argo-Bench,一个包含 210 项数据科学与分析任务的评测框架,基于纽约外卖平台模拟器构建 235 张表、75 亿行的 ERP 仓库(仿 Oracle E-Business Suite schema),要求智能体在仓库中导航重建事实并执行封禁欺诈账户、分配骑手激励预算等动作,由模拟器按后果评分。

正文

Published on Oct 1

Authors:

,

,

,

Abstract

Real-world enterprise data science and analytics workflows require reasoning across dozens of tables, performing statistical analyses, and acting on the results. Established text-to-SQL benchmarks evaluate query generation alone, and audits have found their answer keys frequently wrong. Because real enterprise warehouses are too sensitive to release, these benchmarks are built on public datasets where a business event fits in a single table. We introduce Argo-Bench, an evaluation framework comprising 210 data science and analytics tasks. Drawing on public data, peer-reviewed industry literature, and regulatory filings, we simulate a food delivery platform in New York City at true scale, with 81 million orders in 2024, grounded economics, fraud patterns, and marketplace incentives. We export this world to an ERP warehouse of 235 tables and 7.5 billion rows, modeled on the Oracle E-Business Suite schema. The simulator's ground-truth state is withheld from the warehouse the agent sees, so tasks require reconstructing facts by navigating the warehouse before acting on them. Argo-Bench goes beyond text-to-SQL: the agent files actions such as banning fraudulent accounts, allocating courier incentive budgets, or issuing back pay, and the grader scores each by its consequences in the simulator. Every task has an executable reference solution that demonstrates solvability using only the warehouse. The strongest of 14 frontier and open-weight models scores 95 or higher on only 34.8% of tasks and averages 59.5 points. We hope Argo-Bench drives progress toward agents that understand, navigate, and act within real data environments.

View arXiv page View PDF Project page GitHub Add to collection

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2610.02122 in a model README.md to link it from this page.

Datasets citing this paper 2

textql/Argo-Bench

Viewer •

Updated about 2 hours ago

•

3.91B

•

318

•

2

textql/Decision-Bench

Updated 3 minutes ago

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2610.02122 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.

来源:HuggingFace Daily Papers(社区热门论文) · huggingface.co