跳到正文
arXiv:cs.LG· Sasha Voitovych, Adam Block, Alexander Rakhlin, Abhishek Shetty·· 4 小时前AI 评分31

无需采样访问与完美标签:Gaussian FTPL 实现 Oracle 高效无参数 agnostic 平滑在线学习

Oracle-Efficient and Parameter-Free Agnostic Smoothed Online Learning

AI 导读

研究者提出首个在 agnostic 设定下无需知晓基础测度 μ 的 oracle 高效算法,基于 Gaussian Follow-The-Perturbed-Leader,无需 μ、平滑参数 σ 或时间跨度 T 的先验知识。该算法对 VC 维为 d 的二分类问题取得 Õ(d√(T/σ)) 的 regret,每轮仅调用一次 ERM oracle,结果在 √d 因子内达到最优。

正文

View PDF HTML (experimental)

Abstract:Online learning is an attractive framework in many domains because it permits well-defined learning even when data are dependent or chosen adversarially. This generality, however, comes at a steep price, introducing significant statistical and computational barriers. Recently, smoothed online learning has emerged as a promising framework that interpolates between the fully adversarial and fully stochastic settings by assuming that the conditional law of each covariate has density at most $1/\sigma$ with respect to some fixed base measure $\mu$, and it is known to match the statistical and computational guarantees of classical learning while still allowing for much of the flexibility of online learning. However, existing oracle-efficient algorithms require either (i) sampling access to the base measure $\mu$ or (ii) labels that are perfectly predicted by a fixed hypothesis. Both assumptions limit the applicability of these algorithms, in contrast to statistical learning, where empirical risk minimization (ERM) learns efficiently in the agnostic setting without any knowledge of the data distribution. We show that neither assumption is necessary, giving the first oracle-efficient algorithm that achieves sublinear regret in the agnostic setting without knowledge of $\mu$. Our algorithm, based on Gaussian Follow-The-Perturbed-Leader, is parameter-free: it requires no knowledge of $\mu$, the smoothing parameter $\sigma$, or the horizon $T$, and it achieves regret $\widetilde O(d\sqrt{T/\sigma})$ for binary classes of VC dimension $d$ with a single call to an ERM oracle per round, which is optimal up to a $\sqrt{d}$ factor. En route to establishing the regret bound, we introduce several new techniques that may be of independent interest.
Subjects: Machine Learning (cs.LG); Machine Learning (stat.ML)
Cite as: arXiv:2610.10499 [cs.LG]
  (or arXiv:2610.10499v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.10499

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Sasha Voitovych [view email]
[v1] Wed, 7 Oct 2026 17:49:17 UTC (36 KB)

来源:arXiv:cs.LG · arxiv.org