arXiv:cs.AI· Yu Li, Guangfeng Cai, Long-Fei Li, Shuo Han, Shengtian Yang, Han Luo, Kaibing Yang, Lei Feng·· 5 小时前AI 评分35
DepGPO:面向终端智能体的依赖感知策略优化
Credit Where It Matters: Dependency-Aware Policy Optimization for Terminal Agents
AI 导读
针对终端智能体强化学习中信用分配未追踪命令读写依赖的问题,研究者提出依赖感知组策略优化(DepGPO),从执行轨迹构建命令依赖图,并从任务验证器检查的资源反向追踪,将信用分配给相关写入及其支撑读取,再据此在步骤间重新分配轨迹优势。对比实验与消融研究表明,DepGPO 在复杂终端任务上提升了任务表现与训练稳定性。
正文
Abstract:Terminal-using agents benefit from reinforcement learning (RL) in coding, debugging, and other multi-step terminal tasks. In these tasks, later commands often depend on information or intermediate results produced by earlier commands. However, existing trajectory-level and step-level credit assignment methods do not explicitly trace the read-write dependencies through which commands affect the final outcome. Consequently, training signals could still be assigned to irrelevant operations, weakening learning from relevant steps. In this paper, we propose Dependency-Aware Group Policy Optimization (DepGPO), which uses execution dependencies between commands to guide credit assignment for terminal agents. Specifically, we construct a command dependency graph from execution traces and trace backward from the resources inspected by the task verifier. We then assign credit to relevant writes and their supporting reads along these paths, and use it to redistribute trajectory advantages across steps. Extensive comparative experiments and ablation studies demonstrate that DepGPO improves task performance and training stability on complex terminal tasks.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.03634 [cs.AI] |
| (or arXiv:2610.03634v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03634 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yu Li [view email]
[v1]
Fri, 2 Oct 2026 17:24:48 UTC (696 KB)
来源:arXiv:cs.AI · arxiv.org