AutoCompact 训练编码智能体自行决定何时压缩上下文、保留哪些工作状态以及如何恢复,在 SWE-bench Verified 上通过率提升 9.2 分,SWE-PolyBench Verified 提升 5.0 分。
First AutoHarness, then AutoContext, now AutoCompact.
I am seeing a rising trend of work that trains models to natively support more of what the harness does.
This work specifically trains agents to decide for themselves when to compact.
Reminds me of the new paper from Meta that trains models to manage context natively.
But how good is this approach?
AutoCompact trains the agent to decide when to compact, what working state to keep, and how to resume.
A judge first reviews the base agent's compaction decisions and replaces flawed ones before they execute. The corrected trajectories are used for SFT, then RL with task-success rewards trains coding and compaction together.
Pass rates improve by 9.2 points on SWE-bench Verified and 5.0 points on SWE-PolyBench Verified.
The gain holds even with a 256K window that never overflows, so learned compaction helps when context space is not the limit.
It remains to be seen how this works at scale and how robust it is across harnesses.
One interesting note from the authors is that this type of proactive compaction is a form of model-harness co-design: the harness provides the compaction mechanism, while the model learns when to invoke it, what to preserve, and how to continue afterward.
Even more interesting is how to combine the rule-based compaction techniques already packaged in harnesses with more model-invoked proactive ones.
Paper: https://arxiv.org/abs/2610.02163
Chat with Paper: https://academy.dair.ai/papers/autocompact-learning-when-to-compact-context-in-long-horizon-coding-agents-2610.02163
来源:elvis · x.com