跳到正文
arXiv:cs.AI· Haoyu Zhao, Zhengxu Yu, Zhiyuan He, Meng Fang, Rasul Tutunov, Haitham Bou-Ammar, Weilin Luo, Jun Wang·· 3 小时前

Memento 3:通过反思式规则手册实现基于模型的递归自我改进

Memento 3: Model-Based Recursive Self-Improvement through Reflective Rulebooks

AI 导读

Memento 3 让冻结的 LLM 智能体通过外部记忆持续学习显式世界模型:以自然语言规则手册作为持久语义记忆,并经编译、验证的代码执行预测与规划。在 ARC-AGI-3 上,单模型智能体通关全部 25 个公开游戏的所有关卡,平均 RHAE 达 100.0,仅用人类 44% 的动作数。在 Atari Pong 案例中,学得的反馈控制器在三局不同开局中均以 21:0 获胜,且无需再调用 LLM。

正文

View PDF HTML (experimental)

Abstract:Learning to act in unfamiliar environments requires agents to infer how the world works and revise that understanding as new evidence arrives. Yet limited observations can support multiple world models that explain past interactions but predict different outcomes in unseen states. We introduce Memento 3, building on the Memento series to enable frozen LLM agents to continually learn explicit world models through external memory. The agent maintains a natural-language rulebook as persistent semantic memory, recording revisable hypotheses about environment dynamics while leaving unknown aspects underspecified. It compiles this rulebook into executable code for prediction and planning. Through a continual loop of observation, reflection, rule revision, compilation, and verification, the agent uses prediction errors to refine both the rulebook and its code. Updated code is accepted only when the LLM judges it faithful to the rulebook and cell-exact replay reproduces the observed transitions. We investigate this process as a model-based route to recursive self-improvement (RSI): the agent autonomously explores the environment, revises its world model, and uses verified updates to guide subsequent interaction and learning, while the underlying LLM remains fixed. A population extension maintains multiple world models in parallel, sharing interaction evidence and using their predictions to guide exploration. On ARC-AGI-3, the single-model agent clears every level of all 25 public games, achieves a mean Relative Human Action Efficiency (RHAE) of 100.0, and uses 44% of the human action count. In an Atari Pong case study, a learned feedback controller wins 21:0 in each of three evaluated episodes with different openings, without further LLM calls.
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
Cite as: arXiv:2610.11794 [cs.AI]
  (or arXiv:2610.11794v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.11794

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Haoyu Zhao [view email]
[v1] Thu, 8 Oct 2026 12:06:54 UTC (2,034 KB)

来源:arXiv:cs.AI · arxiv.org