arXiv:cs.AI· Yuyang Dai, Yushun Dong·· 6 小时前AI 评分38
轨迹检索投机解码:模型自身历史何时有用?
Trajectory-Retrieval Speculative Decoding: When Does a Model's Own History Help?
AI 导读
研究提出 Trajectory-Local Adaptive Retrieval(TLAR),从当前推理轨迹中检索近似匹配的续写内容,并结合近期验证结果自适应调整检索激活与候选宽度,与模型生成的草稿共享候选树,经精确验证保持目标模型输出分布。在代码调试、数学与开放式写作任务中,TLAR 与强检索基线结合,在相同验证预算下提升 token 接受率,端到端吞吐量超过草稿模型基线。
正文
Abstract:Long chain-of-thought reasoning increases sequential decoding cost while creating a growing history of potentially reusable continuations. We investigate when this history supplies useful drafts and complements an existing drafter. Controlled source comparisons reveal trajectory-specific reuse, motivating our method Trajectory-Local Adaptive Retrieval (TLAR). TLAR retrieves approximately matched continuations from the current trajectory and uses recent verification outcomes to adapt retrieval activation and candidate width. TLAR combines retrieved continuations with model-generated drafts in a shared candidate tree, preserving the target model's output distribution through exact verification. Across code debugging, mathematics, and open-ended writing, our evaluation connects source reuse, incremental acceptance, and execution cost. Combining TLAR with strong retrieval baselines improves token acceptance under matched verification budgets and increases end-to-end throughput over the draft-model baseline. These findings support generated trajectories as runtime memory for adaptive inference.
| Comments: | 33 pages |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.07350 [cs.AI] |
| (or arXiv:2610.07350v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07350 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yuyang Tsai [view email]
[v1]
Mon, 5 Oct 2026 20:19:28 UTC (402 KB)
来源:arXiv:cs.AI · arxiv.org