跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Bo-Wen Zhang, Junwei He, Maoqi Liu, Feiran Li, Song-Lin Lv, Wentao Ma, Rongyi Lin, Shuhan Zhong, Lan-Zhe Guo·· 19 小时前AI 评分36

T2SPO:面向智能体强化学习的轨迹到步骤策略优化

T2SPO: Trajectory-to-Step Policy Optimization for Agentic Reinforcement Learning

AI 导读

T2SPO 通过成功交互轨迹为策略学习提供步骤级反馈,用预训练 TabPFN 回归器估计每个状态到成功的剩余距离,相邻状态间的距离变化作为辅助信用信号。在 ALFWorld 和 WebShop 上,1.5B 与 7B 语言模型的整体任务成功率均持续优于 GRPO。

正文

View PDF HTML (experimental)

Abstract:Reinforcement learning enables large language model (LLM) agents to learn multi-step behaviors through interaction with their environments. However, rewards in many interactive tasks reflect only the final outcome, providing limited guidance on which intermediate decisions advance the task. Successful training trajectories contain intermediate states that can provide supervision for subsequent interactions. We introduce Trajectory-to-Step Policy Optimization (T2SPO), a method that uses past interaction trajectories to provide step-level feedback for policy learning. T2SPO derives remaining-distance targets from successful trajectories and pairs them with representations of the states visited along the way. Conditioned on these examples, a pretrained TabPFN regressor estimates the remaining distance to success at each state of a new rollout. Changes in this distance estimate across consecutive states yield auxiliary credit for agent steps alongside task-level supervision. As training proceeds, newly completed trajectories refresh the estimator's context, incorporating new experience without updating its parameters. Experiments with 1.5B and 7B language models on ALFWorld and WebShop show that T2SPO consistently improves overall task success over GRPO.
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.00388 [cs.LG]
  (or arXiv:2610.00388v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.00388

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Bowen Zhang [view email]
[v1] Wed, 30 Sep 2026 10:34:27 UTC (216 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org