Microsoft 与合作者发布 ScholarEvolve 论文,提出基于已发表 agent 研究而非智能体自身失败日志来演化 agent harness。框架将 harness 拆分为工具使用、记忆管理和任务执行模块,对近期论文做主题建模以识别各模块的改进策略,再实现并测试组合,且可随时间纳入新论文。
New paper from Microsoft and colleagues on evolving agent harnesses.
It's a really cool idea to evolve a harness from published research. Something I have also been testing for the past couple of months.
ScholarEvolve proposes harness changes based on published agent research rather than the agent's own failure logs.
It splits the harness into modules for tool use, memory management, and task execution.
It runs topic modeling over recent papers to identify distinct improvement strategies for each module, then implements and tests combinations.
New papers can be added over time.
With the model held fixed, Qwen3.5-27B goal completion on AppWorld Challenge rises from 49.6% to 63.6%, and GPT-5.4-mini on Tau2-Bench Telecom rises from 72.7% to 81.9%.
Paper: https://arxiv.org/abs/2609.40169
Chat with Paper: https://academy.dair.ai/papers/learning-from-research-toward-lifelong-agent-harness-evolution-2609.40169
来源:elvis · x.com