跳到正文
arXiv:cs.LG· Zhen Xu, Qizheng Zhang, Gerry Wan, Shang Zhu, Ce Zhang·· 3 小时前AI 评分43

LEAP:为 LLM 智能体学习高效动作提议

LEAP: Learning Efficient Action Proposals For LLM Agents

AI 导读

针对 LLM 智能体逐步执行导致 rollout 缓慢的问题,研究者提出 LEAP,通过用目标动作序列训练小型 drafter 来提升动作推测的准确率。仅用 0.6B 模型,LEAP 在多数决策上与目标模型一致,使智能体端到端 wall clock 时间最多加快 60%,且任务成功率无系统性变化。该 drafter 还支持无需预先收集轨迹的在线训练,效果与离线训练相当。

正文

View PDF HTML (experimental)

Abstract:LLM agents are known to be slow in rollouts. An agent completes a task one step at a time. At each step, it reasons and then chooses an action to execute. The next step and action cannot start until the previous one has finished. Speculative decoding accelerates the rollouts at the reason phase by drafting and verifying the inference tokens. Recent works have also started to apply similar ideas at the action phase. These works use off-the-shelf models, usually large, to draft action proposals for target model to verify. Large drafters match the target more often but take longer to propose, while small off-the-shelf models are fast but rarely make the same decision as the target. We ask a more general question: what determines the end-to-end speedup of action speculation? To answer it, we develop a latency framework for the speculative round. The framework compares what a round gains with what it costs. The gain depends on how well the drafter predicts the target and on how many steps the task can take before it ends. The cost comes from drafting, from waiting for target verification and from executing tools. Guided by the framework, we introduce LEAP (Learning Efficient Action Proposals) which keeps the drafter small and makes it accurate by training it on the target actions sequences. With a small 0.6B model, LEAP agrees with the target on most decisions and makes agents up to 60% faster in end-to-end wall clock time, with no systematic change in task success. Across various datasets, target models and draft models, the framework accounts for most of the measured speedups. We also show the draft model can be online trained with no prior trace collection and match the performance of offline training, making LEAP practical to deploy in the real world.
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as: arXiv:2610.02670 [cs.LG]
  (or arXiv:2610.02670v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.02670

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Zhen Xu [view email]
[v1] Fri, 2 Oct 2026 01:44:03 UTC (176 KB)

来源:arXiv:cs.LG · arxiv.org