arXiv:cs.LG· Lan Shi, Daigo Shishika, Xuan Wang·· 4 小时前AI 评分30
通过有限深度策略敏感性适应智能体行为变化
Adapting to Changes in Agent Behavior via Finite-Depth Policy Sensitivity
AI 导读
研究者提出一种有限深度框架,通过近似策略 Hessian 与混合导数来估计策略敏感性,从而适应其他智能体行为变化。该方法引入可调传播深度以截断轨迹上的导数传播,并推导出截断误差界,证明误差随传播深度增加而非增、在全时域传播时消失。在以信念驱动的追逃博弈验证中,传播深度越大导数估计误差越低,该方法在估计精度与策略适应上均优于基线,基于敏感性的初始化在零样本回报和后续微调上均优于直接迁移。
正文
Abstract:Adapting a reinforcement learning policy to changes in another agent's behavior typically requires a large amount of new interaction data. Policy sensitivity provides a first-order prediction of how a locally optimal policy changes with a behavioral parameter, but its computation requires second-order derivatives whose effects propagate across future interactions. We develop a finite-depth framework to estimate this sensitivity by approximating the policy Hessian and mixed derivative using information from a reference environment. The method features an adjustable propagation depth which determines where derivative propagation along the trajectory is truncated. We characterize the derivative contributions omitted by finite-depth propagation and derive truncation-error bounds for the approximated derivatives and resulting policy sensitivity. The bounds are nonincreasing with propagation depth and vanish at full-horizon propagation. Using a belief-driven pursuit-evasion game as a validation scenario, the proposed method generally achieves lower derivative-estimation errors as the propagation depth increases and outperforms the baseline methods in both estimation accuracy and policy adaptation. The sensitivity-based initialization improves zero-shot return over direct transfer, and also shows advantages for the subsequent fine-tuning in the target environment.
| Comments: | 8 pages, 4 figures |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.07475 [cs.LG] |
| (or arXiv:2610.07475v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07475 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Lan Shi [view email]
[v1]
Mon, 5 Oct 2026 22:40:40 UTC (485 KB)
来源:arXiv:cs.LG · arxiv.org