跳到正文
arXiv:cs.LG· Yupeng Su, Jiayi Tian, Zheng Zhang, Souvik Kundu·· 7 小时前AI 评分45

ReFold:面向长程智能体的免训练可逆轮间上下文折叠

ReFold: Training-Free Reversible Inter-Turn Context Folding for Long-Horizon Agents

AI 导读

arXiv 论文提出 ReFold,一种免训练的渲染层方案,在不改动底层交互历史的前提下压缩模型实际看到的上下文,无需辅助预测器即可去除两类轮间冗余。在五个长程基准和两个前沿 LLM 上,ReFold 最多降低 2.5 倍 token 消耗、使每会话 KV-cache 内存减半且不损失任务成功率,上下文受限下最多避免 92% 的强制压缩。

正文

View PDF HTML (experimental)

Abstract:Long-horizon LLM agents act on an append-only interaction history that is re-sent to the model at every step, so the context and its cost grow with steps until the sessions exceed the context window. Existing methods manage the context through context requirement prediction, relying on additional model calls, heuristic rules, or trained policies. However, these predictive approaches introduce runtime overhead, invalidate prefix caches, and permanently discard content with no guarantee of recovery. To overcome these limitations, we introduce ReFold: a training-free rendering layer that preserves the underlying interaction history while compressing only the model's rendered context. It removes two kinds of inter-turn redundancy without an auxiliary predictor: content an earlier turn already displayed, replaced by a stub, and turns the agent itself reports finished, folded into a one-line note. Both operators use chunked rendering, rewriting the cached prefix once every few steps rather than at every step. Every removal is strictly reversible, a wrong removal costs one restore from the history rather than permanent content loss. Because it operates at the rendering layer, ReFold is plug-and-play across standard ReAct-style harnesses. Evaluations across five long-horizon benchmarks and two frontier LLMs demonstrate that ReFold reduces token consumption by up to 2.5x and halves the KV-cache memory per session without degrading task success rates. Under capped context budgets, it avoids up to 92% of forced compactions. Under concurrent serving workloads, it reduces request queuing delays by up to 100%, accelerating inference by up to 1.7x, while cutting inference costs by up to 3.4x.
Comments: 27 pages, 6 figures, 14 tables
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2610.07863 [cs.CL]
  (or arXiv:2610.07863v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.07863

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yupeng Su [view email]
[v1] Tue, 6 Oct 2026 07:09:54 UTC (258 KB)

来源:arXiv:cs.LG · arxiv.org