arXiv:cs.LG· Peng Xia, Jianwen Chen, Hanyang Wang, Jiaqi Liu, Kaide Zeng, Yu Wang, Siwei Han, Yiyang Zhou, Xujiang Zhao, Haifeng Chen, Zeyu Zheng, Cihang Xie, Huaxiu Yao·· 4 小时前AI 评分45
SkillRL:通过递归技能增强强化学习进化智能体
SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning
AI 导读
研究者提出 SkillRL 框架,通过自动技能发现与递归进化,让 LLM 智能体从原始经验中提炼可复用的高层行为模式。该框架包含分层技能库 SkillBank、自适应检索策略和与策略协同进化的递归机制,在 ALFWorld、WebShop 及七个搜索增强任务上超越强基线 15.3%,并随任务复杂度提升保持稳健,代码已开源。
正文
Authors:Peng Xia, Jianwen Chen, Hanyang Wang, Jiaqi Liu, Kaide Zeng, Yu Wang, Siwei Han, Yiyang Zhou, Xujiang Zhao, Haifeng Chen, Zeyu Zheng, Cihang Xie, Huaxiu Yao
Abstract:Large Language Model (LLM) agents have shown stunning results in complex tasks, yet they often operate in isolation, failing to learn from past experiences. Existing memory-based methods primarily store raw trajectories, which are often redundant and noise-heavy. This prevents agents from extracting high-level, reusable behavioral patterns that are essential for generalization. In this paper, we propose SkillRL, a framework that bridges the gap between raw experience and policy improvement through automatic skill discovery and recursive evolution. Our approach introduces an experience-based distillation mechanism to build a hierarchical skill library SkillBank, an adaptive retrieval strategy for general and task-specific heuristics, and a recursive evolution mechanism that allows the skill library to co-evolve with the agent's policy during reinforcement learning. These innovations significantly reduce the token footprint while enhancing reasoning utility. Experimental results on ALFWorld, WebShop and seven search-augmented tasks demonstrate that SkillRL achieves state-of-the-art performance, outperforming strong baselines over 15.3% and maintaining robustness as task complexity increases. Code is available at this this https URL.
| Comments: | NeurIPS 2026 |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2602.08234 [cs.LG] |
| (or arXiv:2602.08234v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2602.08234 arXiv-issued DOI via DataCite |
Submission history
From: Peng Xia [view email]
[v1]
Mon, 9 Feb 2026 03:17:17 UTC (11,757 KB)
[v2]
Wed, 7 Oct 2026 17:01:42 UTC (11,746 KB)
来源:arXiv:cs.LG · arxiv.org