跳到正文
arXiv:cs.LG· Shangyang Wu, Shuai Zhao, Ziyue Zhu, Jinyang Wu, Anh Tuan Luu, Haoran Luo·· 3 小时前AI 评分34

SCAD:面向长时程智能体的结构化信用分配与蒸馏

SCAD: Structured Credit Assignment and Distillation for Long-Horizon Agents

AI 导读

针对长时程智能体训练中稀疏终端奖励难以区分中间贡献、在线策略蒸馏随学生生成历史增长而丢失教师指导的问题,研究者提出 SCAD,将交互组织为规划与有界子任务执行,在局部上下文中蒸馏执行过程,并通过跨 rollout 子任务前缀树细化规划信用。在所有评测基准上,SCAD 的宏平均准确率较最强训练基线在文本任务上提升 4.48 个百分点,多模态任务上提升 4.19 个百分点。

正文

View PDF HTML (experimental)

Abstract:Training long-horizon agents to solve complex tasks requires effective supervision over extended interaction sequences. However, sparse terminal rewards obscure intermediate contributions, while on-policy distillation can lose informative teacher guidance as student-generated histories grow. To address this problem, we introduce SCAD, which organizes interactions into planning and bounded subtask execution, distills execution in local contexts, and refines planning credit through cross-rollout subtask prefix trees, with planning receiving full terminal credit and execution receiving positive terminal credit and teacher guidance. Across all evaluated benchmarks, SCAD improves macro-average accuracy over the strongest training baseline by 4.48 percentage points for text tasks and 4.19 points for multimodal tasks. SCAD effectively combines outcome-based credit assignment with teacher-guided distillation to improve planning and execution in long-horizon agents.
Comments: 32 pages
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.03372 [cs.LG]
  (or arXiv:2610.03372v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.03372

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Shangyang Wu [view email]
[v1] Fri, 2 Oct 2026 14:29:42 UTC (1,350 KB)

来源:arXiv:cs.LG · arxiv.org