arXiv:cs.AI· Xuan Liu, Jingbin Qian·· 4 小时前AI 评分42
如何区分智能体 RL 的“到达”与“解决”?checkpoint handoff 评估协议给出答案
Reach or Solve? Deep Diving into Agentic RL Gains with Checkpoint Handoffs
AI 导读
研究者提出 checkpoint handoff 评估协议,将智能体强化学习(RL)的收益拆分为“到达”(Reach)与“解决”(Solve)两项指标,无需重新训练即可解耦二者。
正文
Abstract:Reinforcement learning (RL) is widely used to improve language-model agents, and its gains are usually measured by final task success. However, an agent's earlier actions shape the states in which its later decisions are made, so final task success conflates the ability to reach useful states with the ability to complete the task once there. Comparing agents only on the states each one reaches does not separate the two, since each agent is then scored on states selected by its own actions. To address this conflation, we introduce checkpoint handoff, an evaluation protocol that decouples reaching from completing without retraining. One checkpoint acts as a reacher up to a handoff point, and another continues as the solver from the same replayed history. In detail, (1) Reach measures how often a reacher arrives at states that a replayable environment verifies to be a fixed number of actions from success, and (2) Solve measures how often a solver completes the task from identical cloned copies of those states. Crossing supervised fine-tuning (SFT) and RL checkpoints in both roles across two benchmarks and two independently released training pipelines, we find that the gain from switching the solver from SFT to RL is consistently larger when RL is the reacher, at all three model scales on TravelPlanner and on both ALFWorld splits. Further analyses on ALFWorld show that RL improves both Reach and Solve. The solver gain is larger under an RL reacher because RL reaches solvable states more often, and separately measured Reach and Solve gaps recover most of this difference. Because handoff only requires replaying one checkpoint's history under another, agentic RL evaluations can report arrival and completion alongside final success.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2609.19636 [cs.AI] |
| (or arXiv:2609.19636v2 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2609.19636 arXiv-issued DOI via DataCite |
Submission history
From: Xuan Liu [view email]
[v1]
Thu, 17 Sep 2026 03:27:53 UTC (447 KB)
[v2]
Fri, 2 Oct 2026 08:36:58 UTC (442 KB)
来源:arXiv:cs.AI · arxiv.org