跳到正文
arXiv:cs.LG· Ziyang Cai, Christos Ziakas, Vasilis Kontonis, Tim Pearce, Siddhartha Sen, Akshay Krishnamurthy, Shivam Garg, Dimitris Papailiopoulos·· 4 小时前AI 评分45

Fork-and-Flush:让 Autoresearch 智能体跳出“想法盆地”

Fork-and-Flush: Escaping Idea Basins in Autoresearch Agents

AI 导读

研究者提出 fork-and-flush:把 autoresearch 智能体分叉成多条并行轨迹,各自继承累积工作区但使用全新对话上下文,跑完固定时限后从得分最高的轨迹继续。在 13 个长周期科研与工程任务上,该方法在同等算力下相对单次运行和 best-of-N 基线分别提升 66.0% 和 44.4%(min-max 归一化平均分)。

正文

View PDF HTML (experimental)

Abstract:Autoresearch agents tackle open-ended problems by repeatedly proposing candidate solutions, evaluating them, and using feedback to guide subsequent experiments. We show that independent runs of the same agent on the same task often plateau at substantially different scores, with gaps that persist even after considerable additional compute. Embedding their candidate artifacts by functional similarity provides further evidence that trajectories remain in localized regions of the solution space, which we call idea basins. To help agents escape these basins, we study a simple periodic intervention, fork-and-flush. Our method forks the agent into parallel trajectories, each inheriting the accumulated workspace but starting with a fresh chat context. After running each trajectory for a fixed horizon, the agent continues from the highest-scoring one. Across 13 long-horizon research and engineering tasks, with individual agent runs lasting up to several days, fork-and-flush outperformed the single-run and best-of-N baselines by a relative improvement of 66.0% and 44.4%, respectively, on the min-max normalized average score under an equal compute budget.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.07447 [cs.LG]
  (or arXiv:2610.07447v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.07447

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Christos Ziakas [view email]
[v1] Mon, 5 Oct 2026 21:55:52 UTC (1,685 KB)

来源:arXiv:cs.LG · arxiv.org