跳到正文
arXiv:cs.CL· Xinle Deng, Ruobin Zhong, Hujin Peng, Xiaoben Lu, Yanzhe Wu, Guang Li, Buqiang Xu, Yunzhi Yao, Jizhan Fang, Haoliang Cao, Junjie Guo, Yuan Yuan, Ziqing Ma, Yuanqiang Yu, Rui Hu, Baohua Dong, Hangcheng Zhu, Ningyu Zhang·· 3 小时前

MemTrace:追踪与归因大语言模型记忆系统中的错误

MemTrace: Tracing and Attributing Errors in Large Language Model Memory Systems

AI 导读

研究者提出 MemTrace 框架,将 LLM 记忆流水线转化为可执行的记忆演化图,并构建 MemTraceBench 基准,覆盖 Long-Context、RAG、Mem0、EverMemOS 等代表性记忆系统,用于系统研究记忆失效模式。

正文

Authors:Xinle Deng, Ruobin Zhong, Hujin Peng, Xiaoben Lu, Yanzhe Wu, Guang Li, Buqiang Xu, Yunzhi Yao, Jizhan Fang, Haoliang Cao, Junjie Guo, Yuan Yuan, Ziqing Ma, Yuanqiang Yu, Rui Hu, Baohua Dong, Hangcheng Zhu, Ningyu Zhang

View PDF HTML (experimental)

Abstract:Memory is essential for enabling large language models to support long-horizon reasoning, yet existing memory systems remain unreliable and difficult to debug. Tracing memory's dynamic evolution is crucial to understand how information is synthesized, propagated, or corrupted over time. In this work, we study the new problem of error tracing and attribution in LLM memory systems. We propose a novel framework that transforms memory pipelines into executable memory evolution graphs, enabling fine-grained tracing of operational information flow. We then construct MemTraceBench, a benchmark collected from representative memory systems such as Long-Context, RAG, Mem0, and EverMemOS, to systematically study memory failure modes. We further introduce an automatic attribution method that iteratively traces operation subgraphs to pinpoint the root cause of any failed case. Our analysis reveals that memory failures are systematic, stemming from operation-level issues like information loss and retrieval misalignment. Crucially, we leverage these fine-grained attribution signals to guide downstream prompt optimization, establishing a closed-loop system that automatically corrects faults and boosts end-task performance by up to 7.62 percentage points. Code has ben released at this https URL.
Comments: Ongoing work
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2605.28732 [cs.CL]
  (or arXiv:2605.28732v4 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2605.28732

arXiv-issued DOI via DataCite

Submission history

From: Ningyu Zhang [view email]
[v1] Wed, 27 May 2026 16:53:53 UTC (6,833 KB)
[v2] Sun, 5 Jul 2026 04:42:33 UTC (6,835 KB)
[v3] Thu, 16 Jul 2026 08:06:40 UTC (6,865 KB)
[v4] Thu, 8 Oct 2026 04:00:52 UTC (6,883 KB)

来源:arXiv:cs.CL · arxiv.org