跳到正文
arXiv:cs.LG· Yingying Fan, Yuxuan Han, Jinchi Lv, Xiaocong Xu, Zhengyuan Zhou·· 4 小时前AI 评分29

方差感知 UCB 下的分配稳定性与 Wald 推断

Allocation Stability and Wald Inference under Variance-Aware UCB

AI 导读

研究针对双臂固定时域的方差感知 UCB 策略,给出了由奖励差距与方差决定的临界判据:最优臂拉取次数能否以消失的相对误差做确定性近似,而次优臂次数始终稳定。只要各臂的拉取次数与奖励方差之积依概率发散,臂均值线性组合的普通 Wald 统计量对任意固定非零系数向量都有标准正态极限;但该高斯近似在确定性非零系数向量上一致成立,当且仅当最优臂次数稳定。

正文

View PDF HTML (experimental)

Abstract:Allocation stability is often used to justify Gaussian inference from bandit data, but when is it necessary? In this paper, we address this question for a two-armed, fixed-horizon variance-aware UCB policy with bounded reward distributions that may vary with the horizon. We find a sharp criterion in terms of the reward gap and variances that determines whether the optimal-arm count admits a deterministic approximation with vanishing relative error, while the suboptimal-arm count is always stable. Despite the possible instability of the optimal-arm count, we show that the ordinary Wald statistic for a linear combination of the arm means has a standard normal limit for every fixed nonzero coefficient vector, provided the product of the pull count and reward variance diverges in probability for each arm. Under the same condition, however, this Gaussian approximation holds uniformly over deterministic nonzero coefficient vectors if and only if the optimal-arm count is stable. The analysis relies on two main ingredients: (i) a pathwise comparison with an auxiliary policy whose final optimal-arm count is asymptotically equivalent to the original count and independent of the optimal-arm reward sequence; and (ii) joint limits for the rescaled optimal-arm count and the two studentized sample-mean errors under the original policy, which yield nonstandard Wald limits for certain linear combinations of the arm means with coefficients that vary with the horizon.
Subjects: Machine Learning (stat.ML); Machine Learning (cs.LG); Statistics Theory (math.ST)
Cite as: arXiv:2412.08843 [stat.ML]
  (or arXiv:2412.08843v3 [stat.ML] for this version)
  https://doi.org/10.48550/arXiv.2412.08843

arXiv-issued DOI via DataCite

Submission history

From: Yuxuan Han [view email]
[v1] Thu, 12 Dec 2024 00:44:43 UTC (653 KB)
[v2] Sun, 16 Feb 2025 06:55:38 UTC (74 KB)
[v3] Tue, 6 Oct 2026 18:11:49 UTC (61 KB)

来源:arXiv:cs.LG · arxiv.org