跳到正文
arXiv:cs.LG· Jiangang Han·· 3 小时前AI 评分52

WebFovea 技术报告:视觉 Web Agent 在真实网站上模型判断正确但点击失败的四阶段归因

WebFovea: When the Model Is Right but the Click Is Wrong -- Reliable Round Trips for Vision-Based Web Agents on Live Websites

AI 导读

WebFovea 是一个视觉 Web Agent,在 WebRetriever Challenge 2026 中以 57.0/100 获得第二名,任务是从入口 URL 出发在真实网站上操作并返回可验证答案。

正文

View PDF HTML (experimental)

Abstract:We present WebFovea, a vision-based web agent that placed 2nd in the WebRetriever Challenge 2026 with a final score of 57.0 out of 100. The challenge evaluates agents end to end on Protocol III of the WebRetriever benchmark (arXiv:2607.06118): starting from an entry URL on a live website, the agent must operate the site's own interface and return a verifiable answer. A capable multimodal large language model (LLM) is necessary for this, but not sufficient. The model's decisions reach the browser through the harness, the code between the model and the page. At every step, four things must go right: the model's reply must be parsed into the intended action, the action must take effect on the page, the result must be reported back accurately, and the model must be shown the information it needs. On real websites, many of the failures we observed occurred at one of these four stages rather than in the model's reasoning. A coordinate-space mismatch placed every click at 3/4 of its intended coordinates; actions on native dropdowns, inside iframes, and in text boxes failed silently; and self-generated chat-template tokens contaminated 4.9% of task episodes. WebFovea hardens each stage and surrounds the loop with guardrails that keep the agent within the rules and its budget. The four-stage view does not depend on the model, although some individual fixes do. Because we used the same model in all four submissions, the rise of our official hidden-set score from 31.0 to 57.0 reflects changes to the harness, up to run-to-run variance on live sites. We describe the design, the evidence for each component (including negative results), a failure analysis, the limitations, and a roadmap that includes routing different steps to different models.
Comments: 10 pages, 4 figures, 7 tables. Technical report of the 2nd-place solution in the WebRetriever Challenge 2026
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
Cite as: arXiv:2610.03036 [cs.LG]
  (or arXiv:2610.03036v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.03036

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Jiangang Han [view email]
[v1] Fri, 2 Oct 2026 09:18:28 UTC (208 KB)

来源:arXiv:cs.LG · arxiv.org