跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Yashdeep Chaudhary, Roberto Armellin, Harry Holt·· 9 小时前AI 评分29

RARL:面向多脉冲星际转移的可达性分析强化学习

Reachability-Informed Reinforcement Learning for Multi-Impulse Interplanetary Transfers

AI 导读

研究者提出可达性分析强化学习(RARL),将中间路点选择置于学习决策核心,用局部一阶可达性映射把速度扰动映射为椭球位置集,再由 Lambert 重构生成机动。在 Earth-Mars 基准上,三次独立训练平均机动成本 10.23 km/s,仅比序列凸规划参考高 1.72%;多状态策略在 10,000 次蒙特卡洛出发中全部无冲量上限违规,单状态策略平均可行率仅 6.49%。

正文

View PDF HTML (experimental)

Abstract:Reinforcement learning offers the prospect of a reusable sequential decision-making mechanism for spacecraft trajectory design, motivating policy interfaces that connect learned decisions to the underlying maneuver geometry. This paper develops Reachability Analysis-Informed Reinforcement Learning (RARL) for deterministic multi-impulse interplanetary transfers, placing intermediate waypoint selection at the center of the learned decision process. Local first-order reachability maps bounded velocity perturbations into an ellipsoidal set of next-node positions, within which the policy selects its waypoint. Lambert reconstruction then determines the corresponding maneuver to reach this selected waypoint along a dynamically consistent ballistic arc, coupling learned transfer-geometry selection with classical astrodynamics. A terminal two-impulse reconstruction completes the rendezvous, supported by a linear maneuver-demand assessment used for reward shaping. Numerical studies characterize this interface on a two-body Earth-Mars benchmark. Across three independent training runs, RARL achieves a mean maneuver cost of 10.23 km/s, 1.72% above a validated local sequential convex programming reference. Training over dispersed initial states extends policy reuse across a departure family with fixed target state and transfer duration. Each of the three independently trained multi-state policies completes all 10,000 held-out Monte Carlo departures without impulse-cap violations, compared with a mean feasibility rate of 6.49% for single-state policies. This broader sampled feasibility is accompanied by a 0.61% increase in mean nominal maneuver cost, without further training across departures. These results demonstrate that a reachability-informed decision interface supports benchmark-quality trajectory construction and policy reuse across dispersed departure conditions.
Comments: Preprint. 23 pages, 10 figures
Subjects: Optimization and Control (math.OC); Machine Learning (cs.LG); Systems and Control (eess.SY)
Cite as: arXiv:2610.01344 [math.OC]
  (or arXiv:2610.01344v1 [math.OC] for this version)
  https://doi.org/10.48550/arXiv.2610.01344

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yashdeep Chaudhary [view email]
[v1] Thu, 1 Oct 2026 09:16:52 UTC (8,019 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org