跳到正文
arXiv:cs.CL· Michael Andreev·· 6 小时前AI 评分44

记忆深度与重构上下文宽度:分层检索的受控评测

Memory Depth and Reconstructed Context Width: A Controlled Evaluation of Hierarchical Retrieval

AI 导读

研究基于 EverMemBench 评估了 LLM 长期对话记忆中结构深度(D1-D4)与上下文宽度两个参数的交互作用,将上下文从 1K 扩至 4K token 可使 Accuracy 提升 10.11-17.98 个百分点,而增加深度并未带来单调增益。

正文

View PDF HTML (experimental)

Abstract:Long-term conversational memory is becoming an integral component of modern LLM systems. Proposed architectures group records by topics and events, construct hierarchies and graphs, and connect facts through causal and temporal relations. We experimentally study the interaction between two memory parameters: structural depth and the width of context supplied to the answer model. Using EverMemBench, we evaluate depths D1-D4, core budgets of 1,024/2,048/4,096 tokens, and additional Production and Oracle conditions up to the full archive. Increasing width from 1K to 4K improves Accuracy by 10.11-17.98 percentage points, whereas increasing depth provides no monotonic gain. Beyond 8-16K, Production performance reaches a plateau while tokens per correct answer continue to increase; Oracle preserves quality on full archives of 68-71K tokens. These results motivate further investigation of large, coherent context blocks instead of progressively deeper memory structures.
Comments: 4 pages, 1 figure. Accepted at the PALM Workshop at NeurIPS 2026
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2610.08300 [cs.CL]
  (or arXiv:2610.08300v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.08300

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Michael Andreev [view email]
[v1] Tue, 6 Oct 2026 13:04:43 UTC (35 KB)

来源:arXiv:cs.CL · arxiv.org