arXiv:cs.LG· Mayand Gulati, Kerong Wang, WeiChen Au·· 3 小时前AI 评分32
e-ATS:非平稳环境下 E-Process 授权的 Thompson Sampling 何时可以离开锚点
When May a Bandit Leave Its Anchor? E-Process-Authorized Thompson Sampling under Non-stationarity
AI 导读
论文提出 e-process-authorized Thompson sampling(e-ATS),为每个臂同时维护全历史和折扣 Beta 状态,由 anytime-valid e-process 先授权折扣状态、再用可逆相关性分数控制其影响,未获授权前完全遵循乐观 Thompson sampling(OTS)。
正文
Abstract:Stationarity rewards memory, but after a change the same history can mislead. We ask when forgetting should be permitted. E-process-authorized Thompson sampling (e-ATS) gives each arm full-history and discounted Beta states. An anytime-valid e-process first authorizes the discounted state, then a reversible relevance score controls its influence. Before authorization, e-ATS exactly follows optimistic Thompson sampling (OTS). Under a Beta-Bernoulli prior-predictive stationary model, e-ATS's probability of ever departing from OTS is at most the chosen $\alpha_E$, without fitted thresholds. Relative to e-ATS, removing authorization increased mean normalized dynamic pseudo-regret by $38.4\%$ on the registered suite but reduced it by $7.5\%$ on the literature-derived replay suite. Therefore, evidence controls when adaptation begins, not whether it always helps.
| Comments: | 25 pages, 3 figures. Accepted to the E-Values Workshop at NeurIPS 2026 (poster) |
| Subjects: | Machine Learning (cs.LG); Machine Learning (stat.ML) |
| Cite as: | arXiv:2610.03646 [cs.LG] |
| (or arXiv:2610.03646v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03646 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Kerong Wang [view email]
[v1]
Fri, 2 Oct 2026 17:31:50 UTC (3,489 KB)
来源:arXiv:cs.LG · arxiv.org