arXiv:cs.AI· Siyuan Liu (Fudan University, Meituan Longcat Team), Fan Yu (Fudan University, Meituan Longcat Team), Dongyu Ru (Meituan Longcat Team), Yizhu Liu (Meituan Longcat Team), Yifan Yang (Meituan Longcat Team), Xuezhi Cao (Meituan Longcat Team), Xunliang Cai (Meituan Longcat Team), Yixin Cao (Fudan University)·· 3 小时前
DENSE:将智能体轨迹蒸馏为证据接地的捷径树以实现自我精炼
DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement
AI 导读
针对无奖励信号和正确性标签下难以复用智能体执行经验的问题,研究者提出 DENSE,将轨迹证据组织为嵌套捷径树,合并冗余尝试、识别已完成子任务,把噪声执行轨迹转为结构化可复用的反馈。在 Terminal-Bench 2.1 上,DENSE 在四个智能体模型上取得最高严格通过率,较初始尝试提升 7.12-21.81 个百分点,重试时智能体 token 减少 19.0-43.6%。
正文
Abstract:Online agent deployments accumulate execution trajectories at massive scale and behavioral diversity, for which predefined annotation criteria hardly exist. Extracting useful evidence therefore demands costly manual annotation or verifier signals that fail to scale, leaving valuable evidence buried among redundant, incomplete, and failed executions. This raises a question: without post-execution rewards or correctness labels, how can reusable experience be distilled from the trajectories themselves? To address this challenge, we introduce DENSE (Distilling Evidence from Nested Subtask Executions), which organizes trajectory-derived evidence into nested shortcut trees. By consolidating redundant attempts, identifying resolved subtasks, and retaining useful steps alongside outstanding requirements, DENSE transforms noisy execution traces into structured and reusable task-solving feedback. To evaluate whether such feedback helps agents retry the same task, we design REFIT, which measures success-rate changes between the initial attempt and feedback-guided retries. Among feedback methods without external outcome supervision, DENSE achieves the highest strict pass rate across four agent models on Terminal-Bench 2.1, improving over initial attempts by 7.12-21.81 percentage points with 19.0-43.6% fewer agent tokens on retries. In addition, on hard tasks DENSE consistently outperforms self-reflection in cumulative pass rate across multiple feedback iterations on all four models, demonstrating its strong potential for continual agent self-improvement.
| Comments: | 43 pages, including appendices |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2609.21423 [cs.AI] |
| (or arXiv:2609.21423v5 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2609.21423 arXiv-issued DOI via DataCite |
Submission history
From: Siyuan Liu [view email]
[v1]
Fri, 18 Sep 2026 07:36:21 UTC (1,352 KB)
[v2]
Wed, 23 Sep 2026 07:07:00 UTC (3,474 KB)
[v3]
Thu, 24 Sep 2026 11:10:34 UTC (3,474 KB)
[v4]
Sat, 26 Sep 2026 12:03:24 UTC (2,098 KB)
[v5]
Thu, 8 Oct 2026 09:07:47 UTC (4,101 KB)
来源:arXiv:cs.AI · arxiv.org