跳到正文
arXiv:cs.LG· Ami Tavory, Noam Touitou, Tal Sarig, Frank Cheng, Ido Guy·· 3 小时前AI 评分39

用残差量化表示改进上下文赌博机:RQ 在 13 个数据集中的 11 个上超越基线

Bandits via Additive Quantized Representations

AI 导读

研究者提出残差量化(RQ)作为表示层,用离线训练的 RQ 码本将连续上下文映射为多层离散质心分配,并通过 shadow 机制动态设定,从而让加性赌博机算法在严格有界内存下获得非线性表达能力。

正文

View PDF HTML (experimental)

Abstract:Contextual bandits require balancing nonlinear reward modeling with online efficiency. Tree ensembles and neural methods capture nonlinearities but require periodic retraining and large replay buffers. Linear models update efficiently per observation with O(1) memory, but are fundamentally restricted to linear reward structures. We propose Residual Quantization (RQ) as a representation layer to bridge this gap. An offline-trained RQ codebook maps continuous contexts into discrete centroid assignments across multiple levels, set dynamically through a shadow mechanism. This enables a spectrum of additive bandit algorithms that achieve nonlinear expressivity with strictly bounded memory. Across 13 datasets, RQ variants beat their non-RQ counterparts on 11 of 13 datasets, often by wide margins, while matching doubling-retrain XGBoost and neural baselines using up to 1000 times less memory.
Comments: 40 pages, 16 figures, 12 tables. Accepted at NeurIPS 2026
Subjects: Machine Learning (cs.LG)
MSC classes: 68T05, 68W27, 62L05
ACM classes: I.2.6
Cite as: arXiv:2610.02440 [cs.LG]
  (or arXiv:2610.02440v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.02440

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Ami Tavory [view email]
[v1] Thu, 1 Oct 2026 20:07:10 UTC (351 KB)

来源:arXiv:cs.LG · arxiv.org