arXiv:cs.AI(全量分类)· Juli Huang·· 5 小时前AI 评分42
智能体该记住什么?在有限记忆评测中分离保留与检索
What Should an Agent Remember? Disentangling Retention from Retrieval in Bounded-Memory Evaluation
AI 导读
一项流式召回基准在相同 300 个种子片段上交叉评测保留与选择规则,发现固定访问条件下查询感知选择将必需事实召回率提升 15.5 个百分点(95% CI:12.8 至 18.2),而同时改变历史访问的混合对比所报告的 68.7 个百分点优势中有 53.2 个百分点来自访问本身。
正文
Abstract:A persistent agent must decide both what to retain as information arrives and what to surface once a query appears, yet memory evaluations can confound these decisions by comparing methods that differ in both retention and selection. We build a streaming-recall benchmark crossing retention and selection rules and evaluate every condition on the same 300 seeded episodes. Holding access fixed, query-aware selection improves required-fact recall by 15.5 percentage points (95% CI: 12.8 to 18.2), whereas a mixed comparison that also changes history access reports a 68.7-point advantage, of which 53.2 points are attributable to access. Under bounded retention, query-aware, dense, and oracle selection reach the retention ceiling, and all 319 observed failures in the bounded recency condition are caused by eviction rather than ranking errors. Recall falls to 0% as targets recede sufficiently far into the past. Repeating the evaluation on SQuAD preserves the retention ceiling while showing that dense retrieval can outperform lexical retrieval on natural text. These results show that bounded-memory evaluations should hold access fixed and report retention and selection separately.
| Comments: | Code available in the accompanying repository |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.00366 [cs.AI] |
| (or arXiv:2610.00366v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00366 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Julia Huang [view email]
[v1]
Wed, 30 Sep 2026 06:38:14 UTC (93 KB)
来源:arXiv:cs.AI(全量分类) · arxiv.org