arXiv:cs.LG(机器学习,全量分类)· Yashdeep Chaudhary, Roberto Armellin, Harry Holt·· 9 小时前AI 评分29
RARL:面向多脉冲星际转移的可达性分析强化学习
Reachability-Informed Reinforcement Learning for Multi-Impulse Interplanetary Transfers
AI 导读
研究者提出可达性分析强化学习(RARL),将中间路点选择置于学习决策核心,用局部一阶可达性映射把速度扰动映射为椭球位置集,再由 Lambert 重构生成机动。在 Earth-Mars 基准上,三次独立训练平均机动成本 10.23 km/s,仅比序列凸规划参考高 1.72%;多状态策略在 10,000 次蒙特卡洛出发中全部无冲量上限违规,单状态策略平均可行率仅 6.49%。
正文
Abstract:Reinforcement learning offers the prospect of a reusable sequential decision-making mechanism for spacecraft trajectory design, motivating policy interfaces that connect learned decisions to the underlying maneuver geometry. This paper develops Reachability Analysis-Informed Reinforcement Learning (RARL) for deterministic multi-impulse interplanetary transfers, placing intermediate waypoint selection at the center of the learned decision process. Local first-order reachability maps bounded velocity perturbations into an ellipsoidal set of next-node positions, within which the policy selects its waypoint. Lambert reconstruction then determines the corresponding maneuver to reach this selected waypoint along a dynamically consistent ballistic arc, coupling learned transfer-geometry selection with classical astrodynamics. A terminal two-impulse reconstruction completes the rendezvous, supported by a linear maneuver-demand assessment used for reward shaping. Numerical studies characterize this interface on a two-body Earth-Mars benchmark. Across three independent training runs, RARL achieves a mean maneuver cost of 10.23 km/s, 1.72% above a validated local sequential convex programming reference. Training over dispersed initial states extends policy reuse across a departure family with fixed target state and transfer duration. Each of the three independently trained multi-state policies completes all 10,000 held-out Monte Carlo departures without impulse-cap violations, compared with a mean feasibility rate of 6.49% for single-state policies. This broader sampled feasibility is accompanied by a 0.61% increase in mean nominal maneuver cost, without further training across departures. These results demonstrate that a reachability-informed decision interface supports benchmark-quality trajectory construction and policy reuse across dispersed departure conditions.
| Comments: | Preprint. 23 pages, 10 figures |
| Subjects: | Optimization and Control (math.OC); Machine Learning (cs.LG); Systems and Control (eess.SY) |
| Cite as: | arXiv:2610.01344 [math.OC] |
| (or arXiv:2610.01344v1 [math.OC] for this version) | |
| https://doi.org/10.48550/arXiv.2610.01344 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yashdeep Chaudhary [view email]
[v1]
Thu, 1 Oct 2026 09:16:52 UTC (8,019 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org