跳到正文
arXiv:cs.AI· Yifei Tao, Xinyu Zhong, Henry Hengyuan Zhao, Fanyi Wang, Tengda Guo, Wentao Qiu, Ying Wang, Liujian Tang·· 5 小时前AI 评分44

DyadMem:智能体如何与用户协作的长期记忆基准

DyadMem: A Long-Term Memory Benchmark of How Agents Work with Users

AI 导读

研究者提出 DyadMem,一个面向长期智能体的双域全流程记忆基准,并给出新定义"用户条件化关系型智能体记忆"(URAM)。该基准包含 3,065 个 episode、50,961 个 session 和 61,210 个 QA 实例,覆盖 6 类记忆,同时标注用户侧记忆与 URAM。

正文

View PDF HTML (experimental)

Abstract:Long-term agents must remember not only what is true about a user, but also how a particular agent should work with that user as their shared history evolves. Existing benchmarks primarily supervise user facts and preferences or experience reusable across users, leaving this relationship-specific agent memory implicit. Additionally, most prior works measure the model solely with final-answer QA over long interaction histories, making the assessment still incomplete and unreliable. To this end, we introduce DyadMem with the proposed new definition User-conditioned Relational Agent Memory (URAM). DyadMem jointly annotates user-side memory and URAM along the same multi-session trajectories, resulting in 6 memory categories. To summarize, it includes 3,065 episodes, 50,961 sessions, and 61,210 QA instances, with extensive session-level Capture and Update gold annotations, query-level Recall support, and two QA settings: Gold-Memory and Full-Pipeline. Across 16 open-weight and 4 proprietary models, Gold-Memory QA is consistently strong, yet Full-Pipeline QA drops sharply. Such a gap explicitly supports our fine-grained evaluation design. Additionally, several quantitative results further reveal low Capture recall, incomplete Recall, and unsafe-deletion issues arising from even the frontier LLMs. We further conduct a rigorous experiment to validate the effectiveness of our URAM and observe the positive effects for all 20 models. In summary, DyadMem is a dual-domain, full-pipeline memory benchmark with extensive annotation efforts for advancing the domain's development.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.03020 [cs.AI]
  (or arXiv:2610.03020v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.03020

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yifei Tao Mr. [view email]
[v1] Fri, 2 Oct 2026 08:57:43 UTC (13,251 KB)

来源:arXiv:cs.AI · arxiv.org