arXiv:cs.CL· Tong Zheng, Xidong Wu, Zheng Zhang, Zhankui He, Chaoyi Zhang, Benjamin Coleman, Ruoqiao Wei, Di Bai, Haolin Liu, Rui Liu, Xue Wang, Yue Zhuan, Wang-Cheng Kang, Renkai Xiang, Heng Huang, Xinwu Cheng, Yunsong Guo·· 6 小时前AI 评分40
Dream-RSI:通过演化世界实现递归自我改进
Dream-RSI: Recursive Self-Improvement through Evolving Worlds
AI 导读
Dream-RSI 是一个面向自主 AI 智能体的递归自我改进探索框架,通过轻量编排层将探索显式化、可编程化,并保持底层智能体不变。其核心思路是用历史发现树构建回放模拟器,在此模拟器中"做梦"以获得即时的低成本 off-policy 反馈来评估和优化探索策略,无需重复昂贵的在线评估。
正文
Authors:Tong Zheng, Xidong Wu, Zheng Zhang, Zhankui He, Chaoyi Zhang, Benjamin Coleman, Ruoqiao Wei, Di Bai, Haolin Liu, Rui Liu, Xue Wang, Yue Zhuan, Wang-Cheng Kang, Renkai Xiang, Heng Huang, Xinwu Cheng, Yunsong Guo
Abstract:Recursive self-improvement is becoming essential for autonomous AI agents, whose progress depends on discovering high-value solutions across complex domains. Effective exploration drives this process, yet managing and improving exploration strategies remains a major bottleneck. Current systems face a fundamental dilemma: fixed strategies fail to adapt as search spaces scale, while online policy optimization must navigate vast meta-search spaces under delayed, expensive feedback from long-horizon rollouts. We introduce \textsc{Dream-RSI}, a framework for scalable, recursively self-improving exploration. A lightweight orchestration layer makes exploration explicit and programmable while leaving the underlying base agent unchanged. Our key insight is that accumulated discovery history can act as a replay simulator over the realized search space. By dreaming within this simulator built from historical discovery trees, \textsc{Dream-RSI} obtains immediate, low-cost off-policy feedback to evaluate and refine exploration policies without repeated, expensive online evaluation. The improved policy is then redeployed online to drive further discovery, continuously expanding the simulator pool in a self-improving loop. Across 9 tasks in 4 domains, \textsc{Dream-RSI} achieves competitive quality and improves discovery efficiency in several settings.
| Comments: | 11 pages |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2609.14858 [cs.CL] |
| (or arXiv:2609.14858v2 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2609.14858 arXiv-issued DOI via DataCite |
Submission history
From: Tong Zheng [view email]
[v1]
Mon, 14 Sep 2026 00:10:47 UTC (777 KB)
[v2]
Tue, 6 Oct 2026 09:37:58 UTC (1,013 KB)
来源:arXiv:cs.CL · arxiv.org