跳到正文
arXiv:cs.LG· Michael Tang, Mahmoud Abdelgalil, Jorge I. Poveda·· 4 小时前AI 评分37

欺骗性 Bandit 问题:探索耦合与多智能体学习的脆弱性

The Deceptive Bandit Problem: Exploratory Coupling and the Fragility of Multi-Agent Learning

AI 导读

研究者提出"欺骗性 Bandit 问题",证明对抗智能体可利用另一个智能体探索信号的泄漏信息,通过耦合自身探索动作注入外部性,将学习动态导向新的稳态"欺骗性纳什均衡(DNE)"。

正文

View PDF HTML (experimental)

Abstract:Randomized exploration is central to bandit learning, multi-agent reinforcement learning, and zeroth-order policy search, yet its independence and privacy are usually only treated as technical assumptions. We show that these properties are critical for security purposes and demonstrate how an adversarial agent can exploit privileged information on another agent's exploration. We analyze a deceiver-victim pair in the minimal two-player strongly monotone setting, where a deceptive player obtains leaked signals that are merely correlated with the victim's exploration. We show that, by coupling their own exploratory action with this information, the deceptive player injects an externality that steers the learning dynamics to a new steady state, called the deceptive Nash equilibrium (DNE). We prove that the deceptive bandit learning (DBL) dynamics converge to an arbitrarily small neighborhood of the DNE while retaining optimal convergence rates. Interestingly, our analysis attains these optimal rates while relaxing second-order smoothness conditions from standard bandit optimization literature. We characterize conditions under which deception strictly shifts the steady state and its effect on the deceiver's cost, illustrating the results in a resource-allocation game.
Subjects: Machine Learning (cs.LG); Systems and Control (eess.SY)
Cite as: arXiv:2610.09120 [cs.LG]
  (or arXiv:2610.09120v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.09120

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Michael Tang [view email]
[v1] Tue, 6 Oct 2026 21:12:21 UTC (382 KB)

来源:arXiv:cs.LG · arxiv.org