跳到正文
arXiv:cs.AI· Qianhan Feng, Zhongzhen Huang, Yakun Zhu, Xiaofan Zhang, Qi Dou·· 6 小时前AI 评分34

HMED:面向自改进智能体的后见元经验蒸馏

Learning from Revision Consequences: Hindsight Meta-Experience Distillation for Self-Improving Agents

AI 导读

HMED(Hindsight Meta-Experience Distillation)通过回到修订发生的原始事件、在相同恢复的发现状态下重新执行旧版与修订版 Meta-Skill,从而在共享条件下观察修订带来的实际变化,并将每次对比蒸馏为可复用的 Meta-Experience。

正文

View PDF HTML (experimental)

Abstract:As agents continuously improve by generating and revising Skills, the process that discovers and refines those Skills becomes a learnable object in its own right. Task-Skills directly act on task execution, whereas Meta-Skills govern how agents discover and improve future Skills; their value therefore emerges through the subsequent search processes they induce. Existing approaches improve Meta-Skills from observed raw Skill-search trajectories and branch outcomes. However, branch performance entangles the effects of the initial discovery state and the Meta-Skill revision that generated the search process, making it difficult to characterize what a particular revision actually changed, and pushing updates toward revisions that benefit from favorable states rather than those that improve the process. We introduce HMED (Hindsight Meta-Experience Distillation), a mechanism for constructing Meta-Experience for self-improving agents. HMED revisits the completed event from which a revision originates and re-executes the incumbent and revised Meta-Skills from the same restored discovery state, so that the changes associated with the revision can be observed under a shared condition. Each comparison is distilled into a Meta-Experience, a structured record that can be reused by future updates, so that even revisions that are not ultimately retained still contribute a learning signal. Across three interactive agent benchmarks and both open-source and closed-source models, HMED consistently improves Skill discovery performance over strong baselines, shifting Meta-Skill learning beyond branch outcomes toward the consequences of changing the improvement process.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.07979 [cs.AI]
  (or arXiv:2610.07979v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.07979

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Qianhan Feng [view email]
[v1] Tue, 6 Oct 2026 08:43:43 UTC (3,332 KB)

来源:arXiv:cs.AI · arxiv.org