微软新论文提出让编码智能体直接分析智能体历史日志来写提示词,效果优于试错式调优工具。在相同日志下,该方法在 4 个智能体 benchmark 中的 3 个上击败调优工具 GEPA,每条提示词成本约 $1.60。论文建议先用现有日志让编码智能体生成提示词,试错调优留到最后榨取剩余收益。
New Microsoft paper shows a coding agent that studies your agent's old logs usually writes a better prompt than trial-and-error tuning tools, so start there.
Many prompt-tuning tools tweak the prompt, rerun the agent, and judge each change from just a few runs.
This paper skips all that and hands your agent's saved logs to an ordinary coding agent. It writes code to count what happens across every run, so it catches repeat mistakes that a handful of runs can hide.
Given the same logs, the coding agent's prompts beat the tuning tool GEPA on 3 of 4 agent benchmarks, at about $1.60 per prompt.
Point a coding agent at the logs you already have first, and save trial-and-error tuning for squeezing out the last gains.
来源:Rohan Paul · x.com