arXiv:cs.AI· Jike Zhong, Ritwick Chaudhry, Xuanbai Chen, Tianchen Zhao, Linghan Xu, Yifan Xing, Nishant Sankaran·· 6 小时前AI 评分49
DSV-Mem:面向 MLLM 智能体专业工作流的多模态记忆评测基准
DSV-Mem: Evaluating Multimodal Memory in Professional Workflows for MLLM Agents
AI 导读
研究者推出 DSV-Mem 基准,用于评估 MLLM 智能体的密集有状态视觉记忆,包含专家审核场景和 1000 道题,覆盖 Current State、Past State、Derived State、Change History 与 Conflict/Refusal 五类。
正文
Abstract:Conversational MLLM agents are increasingly expected to assist in professional workflows, from AI research and engineering design to product management and business operations. Yet this capability remains underexplored: existing benchmarks largely focus on informal, everyday interactions and personal-life scenarios featuring photographic natural images, isolated static artifacts, and recall-oriented questions. In contrast, professional scenarios often involve structured, information-heavy artifacts that undergo frequent revisions and authority updates, and compositional queries requiring reconciliation of many artifact versions while tracking state precisely. To address these challenges, we introduce DSV-Mem, a benchmark for evaluating Dense Stateful Visual Memory. DSV-Mem comprises expert-reviewed scenarios and 1,000 questions across five user-oriented categories (Current State, Past State, Derived State, Change History, and Conflict/Refusal). A Hartley-inspired criterion favors questions with broader visual-evidence inspection demands. We also introduce a generation harness that produces evaluation suites by decoupling state-transition synthesis from conversation filling. Evaluation over 27 configurations spanning frontier and open-weight models and memory management methods reveals that the strongest baseline scores below 45% on DSV-Mem. Analysis surfaces findings: 1) multimodality and information density both contribute to difficulty, but state evolution, particularly the number of governing updates, is the dominant tested factor. Raw conversation/haystack length, OCR, and arithmetic are not the primary bottlenecks; 2) models often fail to verify user premises against prior state updates before answering; 3) increased reasoning effort and memory management methods yield limited gains, whereas state-aware designs prove more effective. The benchmark and code will be publicly released.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.08102 [cs.AI] |
| (or arXiv:2610.08102v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08102 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jike Zhong [view email]
[v1]
Tue, 6 Oct 2026 10:30:14 UTC (6,127 KB)
来源:arXiv:cs.AI · arxiv.org