Meta 发布 RankEvolve 论文,提出一个可靠的多智能体自动研究框架,通过可执行协议(EOP)约束 ML 研究的各阶段与门控,把 Claude Code 和 Codex 作为独立节点互相审查和修复改动。
Must-read paper from Meta on reliable auto-research agents.
If you let coding agents run ML experiments, one silent bug like leaked eval data or a disconnected gradient can invalidate hours of training and every iteration built on it.
RankEvolve enforces each research phase and gate through a compiled protocol. It runs Claude Code and Codex as separate nodes that review and repair each other's changes.
At a matched budget, combining the two products raises execution accuracy from 45.8% for the best single product to 62.5%.
Over twelve iterations on the open-source HSTU recommender, it improved NDCG@10 on MovieLens-20M by 4.48% over the published result.
Paper: https://academy.dair.ai/papers/rankevolve-a-reliable-multi-agent-auto-research-harness-for-evolving-ranking-mod-2609.39551
来源:DAIR.AI · x.com