跳到正文
arXiv:cs.AI· Nick Leenders, Roy Lindelauf, Joost van Oijen, Boris Cule·· 3 小时前

基于约束命令条件强化学习与 Bandit 策略选择的实时策略游戏智能体

Constrained Command-Conditioned Reinforcement Learning with Bandit Strategy Selection in Real-Time Strategy Games

AI 导读

研究者提出一种约束命令条件 PPO 执行器,用于 MicroRTS 环境:离散命令指定经济、军队组成、军事姿态与工人策略等战略目标,执行器负责生成单元级动作,Thompson 采样 bandit 作为策略家依据游戏内观测估计对手策略并选择命令元组。

正文

View PDF HTML (experimental)

Abstract:Deep reinforcement learning agents reach strong performance in real-time strategy games but can be brittle against opponents outside their training distribution. Separating strategic command selection from learned unit control allows different strategies to be selected for different opponents while reusing the same execution policy. This requires an executor that can follow different commands and measurable criteria for assessing whether it does so. We introduce a constrained command-conditioned Proximal Policy Optimization (PPO) policy, the executor, for MicroRTS, a real-time strategy environment. Discrete commands specify strategic objectives and behavioral requirements for economy, army composition, military posture, and worker policy over multiple environment steps; the executor determines the unit-level actions used to fulfill them. A Thompson-sampling bandit acts as the strategist, selecting command tuples from an estimate of the opponent's strategy built from in-game observations rather than opponent identity. In a controlled comparison with a flat PPO baseline trained with the same architecture, budget, curriculum and self-play league, the strategist-executor system wins significantly more often against three of the four strongest opponents on a training map, including the two strongest held-out ones (0.55 to 0.97 and 0.01 to 0.34), with no significant difference against the others.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.11663 [cs.AI]
  (or arXiv:2610.11663v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.11663

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: NIck Leenders [view email]
[v1] Thu, 8 Oct 2026 10:37:33 UTC (178 KB)

来源:arXiv:cs.AI · arxiv.org