arXiv:cs.LG· Vikram Kakaria, Anish Kataria, Anany Kotawala·· 4 小时前AI 评分35
用 LLM 导出协方差的相关性 Thompson Sampling:面向组合半臂老虎机的新方法
Structure, Not Belief: Correlated Thompson Sampling from LLM-Derived Covariance in Combinatorial Semi-Bandits
AI 导读
研究提出一种相关性 Thompson Sampling 方法:让 LLM 对臂做一次分组划分,通过 RBF 核转换为正定相关矩阵 Σ,从而用协方差 Σ 采样、但 Beta 后验仍只由真实奖励更新,使 LLM 只影响采样移动方式而非其信念。
正文
Abstract:Combinatorial Thompson sampling (CTS) draws independent posterior samples for every arm, so its exploration dynamics ignore any relation among arms. We study a minimal change to those dynamics: an LLM is queried once for a partition of the arms, the partition becomes a positive-definite correlation matrix $\Sigma$ through an RBF kernel on cluster ranks, and the per-round posterior sample is drawn with covariance $\Sigma$ while the Beta posteriors are updated from real rewards only, so the LLM shapes how the sampler moves, not what it believes. We give a self-contained Bayesian regret bound for the idealized Gaussian sampler whose information gain splits into a $K\log T$ term from the $K$-cluster structure and a ridge term that grows to $d\log T$: the $\sqrt{d/K}$ improvement over independent sampling is a finite-horizon transient, exact only as the within-cluster correlation tends to one. The correlated sampler reduces regret by 19% over CTS on 16 synthetic Bernoulli families at $T=2{,}500$ (6-7% at $T=25{,}000$ with data-adaptive kernels) and by 41% on the Microsoft MIND-small news benchmark ($d=200$ real articles), while pseudo-observation warm starts give nothing. An LLM-free ablation with a simulated oracle of controlled quality shows that on unstructured instances the gain is a property of the kernel shape (a random partition, or a plain tempering of the sampling noise, reproduces it), while belief injection at matched oracle quality never helps.
| Comments: | 15 pages. Accepted (poster) at DynaFront 2026: Dynamics at the Frontiers of Optimization, Sampling, and Games, NeurIPS 2026 Workshop |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Machine Learning (stat.ML) |
| Cite as: | arXiv:2610.07470 [cs.LG] |
| (or arXiv:2610.07470v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07470 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Vikram Kakaria [view email]
[v1]
Mon, 5 Oct 2026 22:35:58 UTC (203 KB)
来源:arXiv:cs.LG · arxiv.org