arXiv:cs.LG· Nicol\`o Cesa-Bianchi, Matteo Papini·· 4 小时前AI 评分30
m-Set 对抗式 Bandits 与 Winner 反馈
m-Set Adversarial Bandits with Winner Feedback
AI 导读
该研究给出了 m-set 对抗式 bandits 在不同效用(winner 奖励或奖励之和)与反馈模型(winner 索引、winner 奖励、奖励之和及其组合)下的 regret 上下界,并与组合 bandits 和 MNL bandits 的标准界进行对比,揭示设定细微变化对学习率的显著影响。主要技术贡献是 regret 的信息论下界,合成数据实验验证了理论分析。
正文
Abstract:We show upper and lower bounds on the regret of $m$-set adversarial bandits for different utilities (winner reward or sum of rewards) and feedback models (winner index, winner reward, sum of rewards, and their combinations). By comparing to standard bounds for combinatorial and MNL bandits, our results reveal how subtle changes in the setting can have a dramatic impact on the learning rates. Our main technical contributions are the information-theoretic lower bounds on the regret. Experiments on synthetic data confirm our theoretical analyses.
| Subjects: | Machine Learning (cs.LG); Machine Learning (stat.ML) |
| Cite as: | arXiv:2610.10128 [cs.LG] |
| (or arXiv:2610.10128v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10128 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Matteo Papini [view email]
[v1]
Wed, 7 Oct 2026 14:10:29 UTC (1,221 KB)
来源:arXiv:cs.LG · arxiv.org