跳到正文
elvis· @omarsar0 · X·· 2 小时前AI 评分58
AI 导读

NVIDIA 等机构发布 VERA 论文,将基准轨迹转化为 9000 多个可重启的沙箱环境,用 rubric 评分并只保留可运行、可从可观察证据打分的环境,随后同时更新模型权重和 harness。

正文

Banger paper from NVIDIA.

One exciting trend I am seeing is building verifiable environments for your agents and training the harness alongside the model.

Co-evolving the harness and the model is a big part of owning the intelligence stack. And many frontier AI companies have started doing that.

This paper shows how this works:

VERA turns benchmark trajectories into more than 9,000 restartable sandboxes with rubric scoring, and keeps only environments that run and can be scored from observable evidence.

It then updates both the model weights and the harness.

A harness edit is kept only if it passes self-tests and adds at least 5 points on the development set.

The system rejects a model checkpoint if its score drops by more than 20%.

At 9B, the co-evolved agent beats the strongest single-axis baseline by 10.3 points on AutoCoWorkBench and 13.0 points on AutoMedBench. At 27B, it scores 71.6 on AutoCoWorkBench, above Claude Opus 4.8.

The environment corpus is open-sourced.

Paper: https://arxiv.org/abs/2610.05923

Chat with Paper: https://academy.dair.ai/papers/vera-scaling-verifiable-environments-for-agentic-co-evolution-2610.05923

来源:elvis · x.com