跳到正文
arXiv:cs.CL· Tong Zheng, Xidong Wu, Zheng Zhang, Zhankui He, Chaoyi Zhang, Benjamin Coleman, Ruoqiao Wei, Di Bai, Haolin Liu, Rui Liu, Xue Wang, Yue Zhuan, Wang-Cheng Kang, Renkai Xiang, Heng Huang, Xinwu Cheng, Yunsong Guo·· 6 小时前AI 评分40

Dream-RSI:通过演化世界实现递归自我改进

Dream-RSI: Recursive Self-Improvement through Evolving Worlds

AI 导读

Dream-RSI 是一个面向自主 AI 智能体的递归自我改进探索框架,通过轻量编排层将探索显式化、可编程化,并保持底层智能体不变。其核心思路是用历史发现树构建回放模拟器,在此模拟器中"做梦"以获得即时的低成本 off-policy 反馈来评估和优化探索策略,无需重复昂贵的在线评估。

正文

Authors:Tong Zheng, Xidong Wu, Zheng Zhang, Zhankui He, Chaoyi Zhang, Benjamin Coleman, Ruoqiao Wei, Di Bai, Haolin Liu, Rui Liu, Xue Wang, Yue Zhuan, Wang-Cheng Kang, Renkai Xiang, Heng Huang, Xinwu Cheng, Yunsong Guo

View PDF HTML (experimental)

Abstract:Recursive self-improvement is becoming essential for autonomous AI agents, whose progress depends on discovering high-value solutions across complex domains. Effective exploration drives this process, yet managing and improving exploration strategies remains a major bottleneck. Current systems face a fundamental dilemma: fixed strategies fail to adapt as search spaces scale, while online policy optimization must navigate vast meta-search spaces under delayed, expensive feedback from long-horizon rollouts. We introduce \textsc{Dream-RSI}, a framework for scalable, recursively self-improving exploration. A lightweight orchestration layer makes exploration explicit and programmable while leaving the underlying base agent unchanged. Our key insight is that accumulated discovery history can act as a replay simulator over the realized search space. By dreaming within this simulator built from historical discovery trees, \textsc{Dream-RSI} obtains immediate, low-cost off-policy feedback to evaluate and refine exploration policies without repeated, expensive online evaluation. The improved policy is then redeployed online to drive further discovery, continuously expanding the simulator pool in a self-improving loop. Across 9 tasks in 4 domains, \textsc{Dream-RSI} achieves competitive quality and improves discovery efficiency in several settings.
Comments: 11 pages
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2609.14858 [cs.CL]
  (or arXiv:2609.14858v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2609.14858

arXiv-issued DOI via DataCite

Submission history

From: Tong Zheng [view email]
[v1] Mon, 14 Sep 2026 00:10:47 UTC (777 KB)
[v2] Tue, 6 Oct 2026 09:37:58 UTC (1,013 KB)

来源:arXiv:cs.CL · arxiv.org