跳到正文
原文
arXiv:cs.AI(全量分类)· Wensen Wu·· 5 小时前AI 评分68

Kepler 论文提出可审计世界模型,在 ARC-AGI-3 全部 25 个公开游戏取得 100.00 RHAE

Kepler: Auditable World Models for ARC-AGI-3

AI 导读

Wensen Wu 发布开源评测框架 Kepler,将假设表示为可执行的世界模型,并通过回溯转移检查和条件预测检查加以验证。在冻结的 Claude Opus 5 配置下,Kepler 在全部 25 个公开游戏获得服务器验证的 100.00 RHAE,无需逐游戏选模型或按分数重跑;183 个完成关卡中 181 个的最终尝试动作数不超过人类中位数基线。

正文

View PDF HTML (experimental)

Abstract:ARC-AGI-3 evaluates agents in interactive environments whose rules and objectives must be inferred from observation. We present Kepler, an open-source harness that represents hypotheses as executable world models and validates them through retrospective transition checks and conditional prediction checks. Under one frozen Claude Opus 5 configuration, Kepler obtained a server-verified 100.00 RHAE on all 25 public games, with no per-game model selection or score-conditioned reruns. On 181 of 183 completed levels, the final Opus attempt used no more actions than the corresponding median-human baseline. The retained board runs used 8,256 environment actions, of which 7,292 occurred in scored levels. Retained local provider-session records yield 858.0 million tokens, 97.37% cache reads, and a \$777.72 cost at September 1, 2026 API list-equivalent rates. We also report three evaluation failures: source-code leakage that produced an invalid perfect run, agents reconstructing a removed harness in a control condition, and autonomous repair masking a broken planner. A single-game observation case study showed that animation frames contained task-relevant information absent from settled text grids. Across the final Claude Opus 5 and GPT-5.6 Sol boards, 48 of 50 game-model cells reached 100. These results indicate that public-set score alone has limited discriminative value and motivate first-attempt, cost-conditioned, and verification-aware reporting.
Comments: 17 pages. Accepted to the non-archival Interpreting Agent Behavior workshop at NeurIPS 2026. Project: this https URL . Code and public traces available
Subjects: Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
Cite as: arXiv:2610.00834 [cs.AI]
  (or arXiv:2610.00834v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.00834

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Wensen Wu [view email]
[v1] Wed, 30 Sep 2026 23:46:59 UTC (238 KB)

来源:arXiv:cs.AI(全量分类) · arxiv.org