跳到正文
arXiv:cs.AI· Chirag Sharma, Benjamin Fowlersmith, Karime Maamari·· 6 小时前AI 评分48

AgentMemGate:解决对话助手记忆中的推测污染问题

AgentMemGate: Addressing Speculation Contamination in Conversational Assistant Memory

AI 导读

AgentMemGate 是一种写入时门控机制,用于分类提取的陈述是推测、已完成事件、纠正还是其他,将推测保留在记忆之外并附条件管理后续提升或删除。在147个对话的留出集上,Mem0 和 Graphiti 将 35.2% 和 27.3% 的未决计划断言为当前状态;AgentMemGate 在核心基准上将污染从 87.5% 降至零,任务准确率从 65% 升至 95%。

正文

View PDF HTML (experimental)

Abstract:Conversational AI assistants with long-term memory extract facts from user messages into a store consulted in later conversations. A stated plan can enter that store as fact: a user who might move to Seattle may be recorded as already living there. We call this speculation contamination. Final-state memory benchmarks miss this error because they do not probe intermediate state and include few unresolved speculations. We present AgentMemGate, a write-time gate for profile-store memory that classifies extracted statements as speculation, completed event, correction, or other. Speculations remain outside memory, with conditions governing later promotion or deletion. We also contribute a dataset of multi-session conversations in which plans are confirmed, abandoned, or left unresolved. On our 147-conversation held-out set, Mem0 and Graphiti assert unresolved plans as current state for 35.2% and 27.3% of pending plans. On the core benchmark, AgentMemGate eliminates all observed contamination relative to the identical ungated pipeline (87.5% to zero for the most exposed extraction style) and raises task accuracy from 65% to 95%. On the harder held-out set, gated contamination is 3.4% to 5.7% and task accuracy rises by 9 to 13 percentage points. Our analysis identifies field matching as the main remaining bottleneck: realistic speculations often match no profile field and never reach the gate. We release our datasets, prompts, and evaluation code.
Comments: Accepted at the PALM Workshop at NeurIPS 2026. 15 pages, 4 figures
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.07707 [cs.AI]
  (or arXiv:2610.07707v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.07707

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Chirag Sharma [view email]
[v1] Tue, 6 Oct 2026 03:59:41 UTC (29 KB)

来源:arXiv:cs.AI · arxiv.org