微软论文提出 Coding-Agent Skill Distillation(CASD),让一个现成编码智能体对完整智能体轨迹语料做统计分析、找出反复出现的失败模式,并把发现写成一条优化后的提示词,无需环境访问和验证数据。
Banger paper from Microsoft on prompt optimization.
(bookmark it)
The claim that a coding agent reading your logs beats GEPA at prompt optimization
The overall finding is that you want to give a coding agent your full set of agent logs and let it write the analysis code, instead of running a search loop over small batches of trajectories.
CASD has an off-the-shelf coding agent compute statistics over the whole trajectory corpus, find recurring failure modes, read representative episodes and write the findings as rules in one prompt. It needs no environment access and no validation data.
Across ALFWorld, tau2-bench retail and telecom, and Spreadsheet Bench-Verified, one pass improves the unoptimized baseline by 16.6 points on average. GEPA improves it by 10.9 and SkillOpt by 5.3.
Each optimized prompt costs about $1.60, more than 22x cheaper than validation-gated search.
Paper: https://academy.dair.ai/papers/coding-agents-are-strong-prompt-optimizers-2609.26261
来源:DAIR.AI · x.com