跳到正文
arXiv:cs.AI· Jike Zhong, Ritwick Chaudhry, Xuanbai Chen, Tianchen Zhao, Linghan Xu, Yifan Xing, Nishant Sankaran·· 6 小时前AI 评分49

DSV-Mem:面向 MLLM 智能体专业工作流的多模态记忆评测基准

DSV-Mem: Evaluating Multimodal Memory in Professional Workflows for MLLM Agents

AI 导读

研究者推出 DSV-Mem 基准,用于评估 MLLM 智能体的密集有状态视觉记忆,包含专家审核场景和 1000 道题,覆盖 Current State、Past State、Derived State、Change History 与 Conflict/Refusal 五类。

正文

View PDF HTML (experimental)

Abstract:Conversational MLLM agents are increasingly expected to assist in professional workflows, from AI research and engineering design to product management and business operations. Yet this capability remains underexplored: existing benchmarks largely focus on informal, everyday interactions and personal-life scenarios featuring photographic natural images, isolated static artifacts, and recall-oriented questions. In contrast, professional scenarios often involve structured, information-heavy artifacts that undergo frequent revisions and authority updates, and compositional queries requiring reconciliation of many artifact versions while tracking state precisely. To address these challenges, we introduce DSV-Mem, a benchmark for evaluating Dense Stateful Visual Memory. DSV-Mem comprises expert-reviewed scenarios and 1,000 questions across five user-oriented categories (Current State, Past State, Derived State, Change History, and Conflict/Refusal). A Hartley-inspired criterion favors questions with broader visual-evidence inspection demands. We also introduce a generation harness that produces evaluation suites by decoupling state-transition synthesis from conversation filling. Evaluation over 27 configurations spanning frontier and open-weight models and memory management methods reveals that the strongest baseline scores below 45% on DSV-Mem. Analysis surfaces findings: 1) multimodality and information density both contribute to difficulty, but state evolution, particularly the number of governing updates, is the dominant tested factor. Raw conversation/haystack length, OCR, and arithmetic are not the primary bottlenecks; 2) models often fail to verify user premises against prior state updates before answering; 3) increased reasoning effort and memory management methods yield limited gains, whereas state-aware designs prove more effective. The benchmark and code will be publicly released.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.08102 [cs.AI]
  (or arXiv:2610.08102v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.08102

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Jike Zhong [view email]
[v1] Tue, 6 Oct 2026 10:30:14 UTC (6,127 KB)

来源:arXiv:cs.AI · arxiv.org