跳到正文
DAIR.AI· @dair_ai · X·· 3 小时前AI 评分61
AI 导读

Salesforce AI Research 提出 CLIFT,一种训练和测试时扩展 web agent 的共形自验证方法。31B 开源 Gemma-4 web agent 在 9 应用的 WebArena Infinity 上得 74.6%,高于使用 browser use 的 Gemini 3 Flash 的 70.1%,且无需在每步或部署时调用前沿裁判。

正文

Another great paper from Salesforce AI Research.

The finding is that a 31B open Gemma-4 web agent scores 74.6% on the 9-app WebArena Infinity set, above Gemini 3 Flash with browser use at 70.1%.

They got there without calling a frontier judge at every step or at deployment.

CLIFT has the agent answer verification questions about its own rollouts.

A conformal certifier keeps only the questions whose answers agree with a training-time judge, weights them by how much they can be trusted, and adds the result to per-step rewards.

At test time, the same frozen question bank picks between a greedy rollout and a few retries, with no external judge.

The trained agent improves 12.8 points over its base model and wins 7 of 9 apps. The question bank also transfers to GPT-5.5 at test time on VisualWebArena, and a translated bank improves a live-web agent on Online Mind2Web without any training on that benchmark.

Paper: https://arxiv.org/abs/2610.06829

Chat with Paper: https://academy.dair.ai/papers/clift-conformal-self-verification-for-web-agent-training-and-test-time-scaling-2610.06829

来源:DAIR.AI · x.com