arXiv:cs.LG(机器学习,全量分类)· Huancheng Chen, Xiaodi Sun, Zhaoqiong Huang, Shenyang Huang Shreya Singhal, Jingwen Lu·· 14 小时前AI 评分38
SkillSpec:通过表示特化实现共识门控的智能体技能演化
SkillSpec: Consensus-Gated Agent Skill Evolution via Representation Specialization
AI 导读
SkillSpec 是一个两阶段框架,通过共识门控演化与表示特化来优化 LLM 智能体的自然语言技能。共识门控阶段仅在配对评估达成共识时提交更新,要求整体提升充分且每次重复评估的配对增益非负;特化阶段则依据完整优化轨迹中的过程与冗余敏感信号,选择 flat、graph 或 hybrid 结构。在六个基准和三个目标语言模型上,SkillSpec 平均成功率较 SkillOpt 提升 6.89%。
正文
Abstract:Natural-language skills are textual procedural memories through which large language model (LLM) agents retain reusable task knowledge without updating model weights. Existing methods typically treat skills as either static artifacts or monolithic documents optimized using aggregate validation scores as feedback. However, representing a skill as a monolithic document restricts optimization to its textual content, without explicitly modeling the structure through which procedural knowledge is retrieved and executed. We identify a key distinction between learning what knowledge to retain and determining how to organize it: textual updates should first be validated through execution evidence, after which the retained knowledge should be structured according to its procedural dependencies and retrieval requirements. To this end, we introduce SkillSpec, a two-phase framework comprising consensus-gated evolution and representation specialization. In the consensus-gated phase, complementary editing intents generate complete candidate skills. An update is committed only when paired evaluations reach consensus, requiring sufficient overall improvement and non-negative aggregate paired gain in every repeated evaluation. In the specialization phase, signals of process and redundancy sensitivity derived from the full optimization trajectory, including accepted and rejected candidates, guide the selection of a flat, graph, or hybrid this http URL six benchmarks and three target language models, SkillSpec improves average success rate over SkillOpt by 6.89%, averaged across the three models. These results demonstrate that reliable skill evolution and representation specialization address complementary objectives: deciding what knowledge to retain and how to structure it for inference.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.00704 [cs.LG] |
| (or arXiv:2610.00704v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00704 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Huancheng Chen [view email]
[v1]
Wed, 30 Sep 2026 20:52:00 UTC (6,087 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org