arXiv:cs.LG· Chinmay Savadikar, Zhaoyu Zhang, Mingyu Zhao, Shuang Xie, Han Li, Tianfu Wu, Lingyun Wang·· 4 小时前AI 评分41
AMBER:用追加式记忆训练长程 Web 智能体
AMBER: Training Long-Horizon Web Agents through Append-Only Memory
AI 导读
针对长程 Web 智能体交互历史超出上下文预算的问题,研究者提出 AMBER(Append-only Memory Bank for Evidence Retention),让智能体在强化学习下联合学习推理、行动与写入自由形式记忆,追加式规则从结构上保证信息不被删除。
正文
Abstract:Modern language-model agents increasingly interact with external environments over long-horizon, multi-step trajectories, where the accumulated interaction history can quickly exceed practical context budgets. To ensure reliability, agents must maintain factual information over long horizons, remember execution errors and corrective feedback, and track progress across actions. Several approaches have been proposed to achieve this without the need for maintaining the entire execution history in context, such as using the reasoning and action history, learning to maintain a fixed-size memory through an overwrite mechanism, and periodic summarization. Although overwrite memory can in principle retain anything an append-only memory can, it must learn to carry each fact through every subsequent rewrite, which is difficult to learn from sparse outcome rewards; for interactive applications like web agents, we find that trained overwrite memories delete key information required by the trajectory, as well as corrective feedback received from the environment. We introduce AMBER (Append-only Memory Bank for Evidence Retention) - a simple and scalable framework where an agent jointly learns to reason, act, and write free-form memory, while an append-only rule guarantees retention by construction. This allows AMBER to be trained end-to-end with reinforcement learning from outcome rewards without the need for extensive curated SFT data. On WebArena Lite, AMBER improves average success over overwrite-based memory by 4.09 percentage points, increases the fraction of tasks solved in five repeated runs by 4.8 percentage points, and matches an overwrite baseline trained on substantially more expensive curated supervision. AMBER achieves these improvements while maintaining a practical token budget, providing a strong balance between context efficiency, task performance, and reliable long-horizon execution.
| Comments: | 29 pages, 11 figures |
| Subjects: | Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.07118 [cs.AI] |
| (or arXiv:2610.07118v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07118 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Chinmay Savadikar [view email]
[v1]
Mon, 5 Oct 2026 17:11:07 UTC (863 KB)
来源:arXiv:cs.LG · arxiv.org