跳到正文
arXiv:cs.CL· Egor Pakhomov, Erik Nijkamp·· 3 小时前AI 评分42

MemoryAgentBench 冲突消解得分几何?冻结的最后写入解析器给出答案

Bookkeeping, Composition, or Unreachable Gold? Reading MemoryAgentBench's Conflict-Resolution Scores Against a Frozen Last-Write Resolver

AI 导读

研究将 MemoryAgentBench 冲突消解分项按"最新陈述胜出"规则执行为零学习解析器,官方指标下该规则答对 80.25% 题目(三个留出列表上为 74.5%)。

正文

View PDF HTML (experimental)

Abstract:MemoryAgentBench's Conflict Resolution split is read as measuring "selective forgetting". We execute the benchmark's own rule - the newest statement about a fact wins - as a zero-learning resolver frozen on one of the four fact lists. Under the official metric the rule answers 80.25% of the questions (74.5% on the three held-out lists). Of the rest, 67 items have a released gold that the last-write graph cannot reach but overwritten statements would ("The capital of India is New Delhi." superseded by "The capital of India is Grosseto."; gold New Delhi); such items are a third of the multi-hop questions at 262K. Two long-context models and our pre-registered approximate re-implementation of the benchmark's BM25 agent, one retained run per item and outcomes only, score 84.7%, 82.6% and 41.6% on the items the rule solves against 10.4%, 11.9% and 6.0% on those 67. The failures are a reachability split plus a small parser-scope residual; the per-item split, not the aggregate, is the unit at which a score here can be read.
Comments: Accepted at the IAB Workshop (Interpreting Agent Behavior) at NeurIPS 2026 (non-archival). 19 pages
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as: arXiv:2610.09193 [cs.AI]
  (or arXiv:2610.09193v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.09193

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Egor Pakhomov [view email]
[v1] Tue, 6 Oct 2026 22:44:32 UTC (53 KB)

来源:arXiv:cs.CL · arxiv.org