跳到正文
arXiv:cs.AI· Lechen Li (State Key Laboratory of Internet of Things for Smart City, University of Macau, Macau 519000, China, College of Water Conservancy and Hydropower Engineering, Hohai University, Nanjing 210098, China), Rongye Shi (School of Artificial Intelligence, Beihang University, Beijing 100191, China), Wanhuan Zhou (State Key Laboratory of Internet of Things for Smart City, University of Macau, Macau 519000, China)·· 4 小时前

DRL-DCO:用深度强化学习动态编排 DE 与 CMA-ES 的进化搜索算法

Learning to Orchestrate Evolutionary Search: Progression-Aware Deep Reinforcement Learning for Dynamic DE-CMA-ES Coordination in Optimization and Structural Model Updating

AI 导读

研究提出 DRL-DCO 算法,用基于 DDPG 的 actor-critic 智能体统一调控差分进化(DE)与 CMA-ES 的进化搜索,通过渐进感知状态表示和多样性奖励动态分配探索与开发资源,并联动调节种群规模、精英保留与重启机制。

正文

Authors:Lechen Li (1 and 2), Rongye Shi (3), Wanhuan Zhou (1) ((1) State Key Laboratory of Internet of Things for Smart City, University of Macau, Macau 519000, China, (2) College of Water Conservancy and Hydropower Engineering, Hohai University, Nanjing 210098, China, (3) School of Artificial Intelligence, Beihang University, Beijing 100191, China)

View PDF HTML (experimental)

Abstract:Solving high-dimensional structural model updating problems requires an algorithm capable of navigating complex, non-convex landscapes with correlated parameters. Existing hybrid evolutionary algorithms typically rely on static architectures or fixed switching rules, resulting in disjointed search phases. To address this, this study proposes a Deep Reinforcement Learning-governed dynamic DE-CMAES Orchestration (DRL-DCO) algorithm, in which a Deep Deterministic Policy Gradient (DDPG)-based actor-critic agent continuously governs the evolutionary process as a single, unified system rather than a mechanical concatenation of algorithms. Guided by a progression-aware state representation and a diversity-informed reward, the agent fluidly reallocates computational resources between the differencevector-based exploration of Differential Evolution (DE) and the covariance-guided exploitation of CMA-ES, while jointly regulating population size, elite preservation, and a restart mechanism to escape local optima. This allows DRL-DCO to autonomously transition between exploration-dominant, exploitation-dominant, and mixed-strategy regimes across generations. Beyond the training phase, the trained actor can operate in a supervision-free inference mode, where the internalized policy autonomously orchestrates DE and CMA-ES control from observed search states through forward inference alone, without critic evaluation or weight updates, enabling faster deployment while retaining full effectiveness. Validated on high-dimensional single-objective optimization benchmarks and the IASC-ASCE structural health monitoring benchmark, DRL-DCO achieves superior convergence accuracy and robustness compared to state-of-the-art adaptive and hybrid evolutionary algorithms, as well as single-operator DRL-governed baselines.
Subjects: Neural and Evolutionary Computing (cs.NE); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.11546 [cs.NE]
  (or arXiv:2610.11546v1 [cs.NE] for this version)
  https://doi.org/10.48550/arXiv.2610.11546

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Lechen Li [view email]
[v1] Thu, 8 Oct 2026 09:12:13 UTC (2,322 KB)

来源:arXiv:cs.AI · arxiv.org