arXiv:cs.AI(全量分类)· Wensen Wu·· 5 小时前AI 评分68
Kepler 论文提出可审计世界模型,在 ARC-AGI-3 全部 25 个公开游戏取得 100.00 RHAE
Kepler: Auditable World Models for ARC-AGI-3
AI 导读
Wensen Wu 发布开源评测框架 Kepler,将假设表示为可执行的世界模型,并通过回溯转移检查和条件预测检查加以验证。在冻结的 Claude Opus 5 配置下,Kepler 在全部 25 个公开游戏获得服务器验证的 100.00 RHAE,无需逐游戏选模型或按分数重跑;183 个完成关卡中 181 个的最终尝试动作数不超过人类中位数基线。
正文
Abstract:ARC-AGI-3 evaluates agents in interactive environments whose rules and objectives must be inferred from observation. We present Kepler, an open-source harness that represents hypotheses as executable world models and validates them through retrospective transition checks and conditional prediction checks. Under one frozen Claude Opus 5 configuration, Kepler obtained a server-verified 100.00 RHAE on all 25 public games, with no per-game model selection or score-conditioned reruns. On 181 of 183 completed levels, the final Opus attempt used no more actions than the corresponding median-human baseline. The retained board runs used 8,256 environment actions, of which 7,292 occurred in scored levels. Retained local provider-session records yield 858.0 million tokens, 97.37% cache reads, and a \$777.72 cost at September 1, 2026 API list-equivalent rates. We also report three evaluation failures: source-code leakage that produced an invalid perfect run, agents reconstructing a removed harness in a control condition, and autonomous repair masking a broken planner. A single-game observation case study showed that animation frames contained task-relevant information absent from settled text grids. Across the final Claude Opus 5 and GPT-5.6 Sol boards, 48 of 50 game-model cells reached 100. These results indicate that public-set score alone has limited discriminative value and motivate first-attempt, cost-conditioned, and verification-aware reporting.
| Comments: | 17 pages. Accepted to the non-archival Interpreting Agent Behavior workshop at NeurIPS 2026. Project: this https URL . Code and public traces available |
| Subjects: | Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA) |
| Cite as: | arXiv:2610.00834 [cs.AI] |
| (or arXiv:2610.00834v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00834 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Wensen Wu [view email]
[v1]
Wed, 30 Sep 2026 23:46:59 UTC (238 KB)
来源:arXiv:cs.AI(全量分类) · arxiv.org