[Submitted connected 14 Sep 2026]
Authors:Tong Zheng, Xidong Wu, Zheng Zhang, Zhankui He, Chaoyi Zhang, Benjamin Coleman, Ruoqiao Wei, Di Bai, Haolin Liu, Rui Liu, Xue Wang, Yue Zhuan, Wang-Cheng Kang, Renkai Xiang, Heng Huang, Xinwu Cheng, Yunsong Guo
View PDF HTML (experimental)
Abstract:Recursive self-improvement is becoming progressively captious for autonomous AI agents, wherever advancement hinges connected discovering high-value solutions crossed analyzable domains. The driver of this process is effective exploration, however, managing and improving exploration strategies remains a awesome bottleneck. Current systems look a basal dilemma: fixed strategies neglect to accommodate arsenic hunt spaces scale, while online argumentation optimization requires navigating immense meta-search spaces nether delayed and costly feedback complete long-horizon rollouts. We present \textsc{Dream-RSI}, a model for scalable and recursively self-improving exploration. A lightweight orchestration furniture makes exploration definitive and programmable while leaving the underlying coding supplier unchanged. Our cardinal penetration is that accumulated find history tin service arsenic a replay simulator complete the realized hunt space. By performing dreaming successful the replay simulator constructed from humanities find trees, \textsc{Dream-RSI} secures immediate, low-cost off-policy feedback to measure and refine exploration policies without invoking repetitive, costly online evaluations. The improved argumentation is subsequently redeployed online to thrust further discovery, continuously expanding the simulator excavation successful a self-improving loop. Across algorithm engineering, mathematical optimization, and GPU kernel engineering, \textsc{Dream-RSI} achieves competitory aliases improved find value while substantially reducing find costs successful respective settings.Submission history
From: Tong Zheng [view email]
[v1] Mon, 14 Sep 2026 00:10:47 UTC (777 KB)
English (US) ·
Indonesian (ID) ·