arXiv:cs.LG· Yujiang Li, Zhenyu Hou, Yi Jing, Jie Tang, Yuxiao Dong·· 2 天前AI 评分45
CompactionRL:面向长程智能体的上下文压缩强化学习
CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents
AI 导读
CompactionRL 是一种结合上下文压缩的强化学习策略,通过联合优化任务执行与摘要生成,让长程智能体在压缩轨迹上学习。
正文
Abstract:Long-horizon agentic LLMs are increasingly limited by finite context windows, as extended interaction trajectories can exceed the maximum context length before a task is completed. Context compaction offers a natural solution by summarizing previous interaction states and continuing the rollout under a compressed context, but incorporating compaction into reinforcement learning remains underexplored. We propose CompactionRL, a reinforcement learning strategy to train long-horizon agentic LLMs with context compaction. Our approach jointly optimizes task execution and summary generation with token-level loss normalization and cross-segment generalized advantage estimation. This design enables the LLM agents to learn from compacted long-horizon trajectories. We train CompactionRL on top of open models and observe consistent performance gains on agentic coding tasks. CompactionRL enables the open GLM-4.5-Air model (106B-A12B) to achieve Pass@1 scores of 66.4% on SWE-bench Verified and 26.2% on Terminal-Bench 2.0, exceeding the base model under inference-time compaction by 6.6 and 4.9 points, respectively. Built upon GLM-4.7-Flash (30B-A3B), CompactionRL improves Pass@1 by 5.5 and 6.7 points against the base model, reaching 56.0% on SWE-bench Verified and 20.2% on Terminal-Bench 2.0. CompactionRL is thus deployed in the RL pipeline for training the open GLM-5.2 model (750B-A40B).
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2607.05378 [cs.LG] |
| (or arXiv:2607.05378v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2607.05378 arXiv-issued DOI via DataCite |
Submission history
From: Yujiang Li [view email]
[v1]
Mon, 6 Jul 2026 17:55:12 UTC (254 KB)
[v2]
Thu, 1 Oct 2026 09:57:02 UTC (256 KB)
来源:arXiv:cs.LG · arxiv.org