跳到正文
arXiv:cs.LG· Lan Shi, Daigo Shishika, Xuan Wang·· 4 小时前AI 评分30

通过有限深度策略敏感性适应智能体行为变化

Adapting to Changes in Agent Behavior via Finite-Depth Policy Sensitivity

AI 导读

研究者提出一种有限深度框架,通过近似策略 Hessian 与混合导数来估计策略敏感性,从而适应其他智能体行为变化。该方法引入可调传播深度以截断轨迹上的导数传播,并推导出截断误差界,证明误差随传播深度增加而非增、在全时域传播时消失。在以信念驱动的追逃博弈验证中,传播深度越大导数估计误差越低,该方法在估计精度与策略适应上均优于基线,基于敏感性的初始化在零样本回报和后续微调上均优于直接迁移。

正文

View PDF HTML (experimental)

Abstract:Adapting a reinforcement learning policy to changes in another agent's behavior typically requires a large amount of new interaction data. Policy sensitivity provides a first-order prediction of how a locally optimal policy changes with a behavioral parameter, but its computation requires second-order derivatives whose effects propagate across future interactions. We develop a finite-depth framework to estimate this sensitivity by approximating the policy Hessian and mixed derivative using information from a reference environment. The method features an adjustable propagation depth which determines where derivative propagation along the trajectory is truncated. We characterize the derivative contributions omitted by finite-depth propagation and derive truncation-error bounds for the approximated derivatives and resulting policy sensitivity. The bounds are nonincreasing with propagation depth and vanish at full-horizon propagation. Using a belief-driven pursuit-evasion game as a validation scenario, the proposed method generally achieves lower derivative-estimation errors as the propagation depth increases and outperforms the baseline methods in both estimation accuracy and policy adaptation. The sensitivity-based initialization improves zero-shot return over direct transfer, and also shows advantages for the subsequent fine-tuning in the target environment.
Comments: 8 pages, 4 figures
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.07475 [cs.LG]
  (or arXiv:2610.07475v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.07475

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Lan Shi [view email]
[v1] Mon, 5 Oct 2026 22:40:40 UTC (485 KB)

来源:arXiv:cs.LG · arxiv.org