跳到正文
arXiv:cs.LG· Jia Liufu, Bin Hu, Linglin Jing, Terry Kong, Yuki Huang, Ashwath Aithal, Wenming Yang, Jun Yang·· 4 小时前AI 评分38

FC-SWE:面向长程软件工程智能体的失败条件化强化学习框架

FC-SWE: Failure-Conditioned RL for Long-Horizon Software Engineering Agents

AI 导读

FC-SWE 是一种失败条件化强化学习框架,将补丁验证失败后的恢复尝试纳入策略训练:失败后仓库重置至原始任务状态,以失败补丁和 verifier 反馈作为上下文生成恢复轨迹。

正文

View PDF HTML (experimental)

Abstract:Repository-level software engineering (SWE) is a challenging long-horizon setting: agents must reason over extended interactions, use tools, and adapt to stateful environments. Recent work trains SWE agents with reinforcement learning methods such as Group Relative Policy Optimization (GRPO), which independently sample multiple trajectories per issue, test the resulting patches, and compare terminal rewards within a fixed group. However, this training setup does not reuse verifier feedback from failed patches as context for subsequent attempts, even though this feedback contains valuable diagnostic information about what went wrong. Training on recovery trajectories is challenging because the preceding outcome determines whether the next trajectory is generated, while the failed execution determines its conditioning context. We introduce FC-SWE, a failure-conditioned RL framework that incorporates recovery attempts into policy training. After a patch fails verification, FC-SWE restores the repository to its original task state and uses the failed patch and verifier feedback as context for a recovery trajectory. FC-SWE adapts GRPO to these chains of complete, multi-turn tool-use trajectories through two mechanisms. Trajectory-local rewards preserve each attempt's verifier outcome, preventing recovery success from rewarding an earlier failed patch. Active-set advantage estimation forms a comparison group from all initial and recovery trajectories actually executed for the same issue, so failed attempts remain in the group while unexecuted attempts are excluded. On all 500 SWE-bench Verified tasks under a verifier-assisted protocol, FC-SWE with Qwen3.5-4B and SWE-agent achieves 41.7% Resolved@1 and 52.8% Resolved@2, compared with 38.9% and 48.5% for GRPO. Although trained with at most two attempts per chain, FC-SWE reaches 70.7% Resolved@11 under an eleven-attempt test-time budget.
Comments: 23 pages, 8 figures, 6 tables
Subjects: Machine Learning (cs.LG); Software Engineering (cs.SE)
Cite as: arXiv:2610.07898 [cs.LG]
  (or arXiv:2610.07898v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.07898

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Jia Liufu [view email]
[v1] Tue, 6 Oct 2026 07:44:12 UTC (401 KB)

来源:arXiv:cs.LG · arxiv.org