arXiv:cs.LG· Yatai Ji, Zhengqiu Zhu, Yong Zhao, Yue Hu, Fanglong Yao, Chen Gao, Pengfei Zhu, Quanjun Yin·· 4 小时前AI 评分33
SearchWorld:基于世界模型的空间价值想象助力无人机目标搜索
SearchWorld: Spatial Value-Grounded Imagination for UAV Object Search via World Models
AI 导读
SearchWorld 是一个循环状态空间世界模型,将显式空间记忆与价值引导想象结合,用 BEV 探索与障碍记忆解码任务感知的空间价值层来引导搜索,无需单独训练标量 critic。在 UAV-ON 上成功率提升至 23.8%(此前最强已发表智能体为 19.5%),oracle 成功率 35.5%,未见场景下成功率 19.9%。
正文
Abstract:Autonomous unmanned aerial vehicle (UAV) object search involves a closed loop of perception, decision-making, and action under partial observability. Urban environments pose several challenges: large search areas and narrow egocentric views limit coverage, dense 3D geometry constrains safe motion, and open-world instructions require identifying a specific target among distractors. Many existing methods mitigate partial observability through explicit maps or memory representations, yet remain largely reactive, reasoning over past observations without explicitly predicting future states. World models enable prospective reasoning through imagined rollouts. However, image-generating world models can incur high inference latency, while spatially grounded planning remains challenging for latent world models. We propose SearchWorld, a recurrent state-space world model that connects explicit spatial memory with value-guided imagination. The model maintains BEV exploration and obstacle memory and decodes a task-aware spatial value layer to guide search. A cognition-action network uses this learned spatial value prior to improve the policy through imagined rollouts, without training a separate scalar critic. Training progresses from world-model learning to expert imitation and imagination-based exploration refinement. On UAV-ON, SearchWorld improves the success rate to 23.8% (19.5% for the strongest published agent) and raises oracle success to 35.5%, while remaining robust on unseen scenes (19.9% success rate). By grounding imagination in explicit spatial representations, SearchWorld enables UAV agents to plan prospectively rather than react.
| Comments: | 10 pages,2 figures |
| Subjects: | Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.09335 [cs.AI] |
| (or arXiv:2610.09335v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09335 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yatai Ji [view email]
[v1]
Wed, 7 Oct 2026 02:51:02 UTC (966 KB)
来源:arXiv:cs.LG · arxiv.org