NVIDIA 论文提出一种在昂贵的智能体后训练之前评估基座模型潜力的方法。基座模型在 SWE-bench Verified 上几乎无法直接评测,六个基座模型中有五个解出零任务。
Super interesting NVIDIA paper on choosing base models for coding agents.
It's actually a clever way to rank base checkpoints by how well each one is likely to do as a coding agent after post-training.
They document that base models are really hard to evaluate on agentic coding tasks. They ran six base models on SWE-bench Verified, and five of them solved zero tasks.
So instead, they look at the one step in a coding run that actually fixes the task.
In other words, they take tasks that a strong post-trained agent already solved, replay its code edits one by one, and run the tests after each edit. The first edit that makes the tests pass is the decisive edit.
Then they give the base model everything that happened before that edit and check whether it can come up with that fix.
They score this in three ways. They check how likely the base model is to write the fix, whether it can pick the fix out of a set of rejected patches, and whether any fix it writes on its own passes the tests.
All three rankings closely match post-trained SWE-bench Verified scores across ten base and post-trained model pairs.
Why is this useful?
If you pick checkpoints for agentic post-training, this method can give you a signal before you spend the training budget.
来源:DAIR.AI · x.com