跳到正文
arXiv:cs.AI· Hangxi Guo, Fengyuan Liu, Yue Wang, Yuhua Qi, Haoyi Xiong, Fei Sun, Mengnan Du·· 3 小时前

SELF:智能体后见自蒸馏中的环境反馈建模为何重要

Environmental Feedback Modeling Matters: Rethinking Feedback Treatment in Agentic Hindsight Self-Distillation

AI 导读

研究提出 SELF 框架,将环境反馈建模与后见自蒸馏联合优化,让模型在预测环境响应的同时,从反馈条件下的自教师中蒸馏指导。在 Qwen3-8B 上,SELF 在 τ-bench 成功率上分别超过 SDPO 和 GRPO 6.4 和 4.1 个百分点,在 AppWorld 任务目标完成率上分别高出 10.71 和 3.57 个百分点。

正文

View PDF HTML (experimental)

Abstract:Reinforcement learning is commonly used to train language agents in interactive environments, but cannot be directly applied when rewards are unavailable. Recent methods use environmental feedback as privileged context for hindsight self-distillation, but our analysis suggests that simply conditioning the teacher on feedback is insufficient, motivating us to rethink how environmental feedback is used in agentic self-distillation. Given that environmental feedback contains rich supervision for modeling how the environment responds to agent actions, we introduce \textit{agentic SElf-distilLation with environmental Feedback modeling} (SELF), a framework that jointly optimizes environmental feedback modeling and hindsight self-distillation. SELF learns to predict environmental responses while distilling guidance from a feedback-conditioned self-teacher into the policy. Our analysis reveals a mutually reinforcing mechanism: environmental feedback modeling strengthens hindsight supervision and policy learning, while self-distillation enhances the model's ability to model environmental feedback. With Qwen3-8B, SELF outperforms SDPO and GRPO by 6.4 and 4.1 percentage points in $\tau$-bench success rate, and by 10.71 and 3.57 percentage points in AppWorld task goal completion, respectively. These results show that SELF uses environmental feedback more effectively within agentic self-distillation, improving agent capabilities.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.11384 [cs.AI]
  (or arXiv:2610.11384v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.11384

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Fengyuan Liu [view email]
[v1] Thu, 8 Oct 2026 07:16:00 UTC (562 KB)

来源:arXiv:cs.AI · arxiv.org