arXiv:cs.CL· Egor Pakhomov, Erik Nijkamp·· 3 小时前AI 评分42
MemoryAgentBench 冲突消解得分几何?冻结的最后写入解析器给出答案
Bookkeeping, Composition, or Unreachable Gold? Reading MemoryAgentBench's Conflict-Resolution Scores Against a Frozen Last-Write Resolver
AI 导读
研究将 MemoryAgentBench 冲突消解分项按"最新陈述胜出"规则执行为零学习解析器,官方指标下该规则答对 80.25% 题目(三个留出列表上为 74.5%)。
正文
Abstract:MemoryAgentBench's Conflict Resolution split is read as measuring "selective forgetting". We execute the benchmark's own rule - the newest statement about a fact wins - as a zero-learning resolver frozen on one of the four fact lists. Under the official metric the rule answers 80.25% of the questions (74.5% on the three held-out lists). Of the rest, 67 items have a released gold that the last-write graph cannot reach but overwritten statements would ("The capital of India is New Delhi." superseded by "The capital of India is Grosseto."; gold New Delhi); such items are a third of the multi-hop questions at 262K. Two long-context models and our pre-registered approximate re-implementation of the benchmark's BM25 agent, one retained run per item and outcomes only, score 84.7%, 82.6% and 41.6% on the items the rule solves against 10.4%, 11.9% and 6.0% on those 67. The failures are a reachability split plus a small parser-scope residual; the per-item split, not the aggregate, is the unit at which a score here can be read.
| Comments: | Accepted at the IAB Workshop (Interpreting Agent Behavior) at NeurIPS 2026 (non-archival). 19 pages |
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.09193 [cs.AI] |
| (or arXiv:2610.09193v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09193 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Egor Pakhomov [view email]
[v1]
Tue, 6 Oct 2026 22:44:32 UTC (53 KB)
来源:arXiv:cs.CL · arxiv.org