跳到正文
arXiv:cs.AI· Yunqi Liu, Tong Niu, Zitong Wang, Zhenlong Dai, Yuqi Qing, Weiqiang Wang, Jian Liu·· 4 小时前

EgoBench:面向工具使用智能体的交互式第一视角多模态基准

EgoBench: An Interactive Egocentric Multimodal Benchmark for Tool-Using Agents

AI 导读

EgoBench 是首个面向工具使用智能体的交互式多模态基准,包含 1,590 个基于第一视角视频的任务,覆盖五类日常场景。该基准通过多智能体模拟用户和确定性联合验证框架评估智能体的交互能力,八个视频 MLLM 智能体上最佳模型平均 Joint Success Rate 仅 34.95%。该工作已被 NeurIPS 2026 接收。

正文

View PDF HTML (experimental)

Abstract:As AI agents increasingly operate in open, real-world environments, they require a deep synergy of multimodal perception, tool invocation with multi-hop reasoning, and dynamic interaction with users. However, existing benchmarks fail to jointly evaluate these capabilities due to challenges in designing strictly coupled multi-capability tasks, simulating natural and task-constrained user feedback, and ensuring objective evaluation of dynamic interaction. To bridge this gap, we introduce EgoBench, the first interactive multimodal benchmark for tool-using agents. EgoBench comprises 1,590 egocentric-video-grounded tasks covering five daily scenarios, along with a user-agent-tool interactive environment for evaluation. We implement a three-stage synergistic pipeline through which each task is designed to enforce the joint application of visual perception and tool-augmented multi-hop reasoning. We additionally develop a multi-agent simulated user to evaluate agents' interaction capabilities, which generates high-fidelity, task-aligned responses to agents. Furthermore, we establish a deterministic joint validation framework that guarantees objective assessment through process-based and result-based equivalence. Benchmarking eight video-MLLM agents on EgoBench reveals a severe performance ceiling: the best-performing model achieves only a 34.95% average Joint Success Rate across the three interaction modes. Finally, we conduct a multi-dimensional error analysis to disentangle failure modes, exposing capability bottlenecks for advancing future AI agents.
Comments: Accepted to NeurIPS 2026. Camera-ready version
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2605.27820 [cs.AI]
  (or arXiv:2605.27820v2 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2605.27820

arXiv-issued DOI via DataCite

Submission history

From: Tong Niu [view email]
[v1] Wed, 27 May 2026 01:28:15 UTC (10,613 KB)
[v2] Thu, 8 Oct 2026 09:38:28 UTC (8,060 KB)

来源:arXiv:cs.AI · arxiv.org