arXiv:cs.LG· Benedikt Koch, Winston Chou, Aur\'elien Bibaut, Nathan Kallus·· 4 小时前AI 评分37
弱信号下的策略学习:Netflix 大规模实验中的个性化线性平滑策略
Policy Learning with Weak Signals
AI 导读
研究将数字实验中弱信噪比、高维协变量与海量数据下的处理效应估计形式化为有界信噪比的高斯观测,证明一般情况下最优处理策略不可学习、学习最优策略价值的速度也慢得不切实际。当处理效应平滑变化时,基于线性平滑器的极小极大自适应策略可实现趋零的福利遗憾。在 Netflix 大规模真实实验中,个性化线性平滑策略的表现优于非个性化策略。
正文
Abstract:Policy learning in digital experimentation faces three challenges: weak signal-to-noise ratios, rich covariate spaces, and massive data volumes. We formalize this regime by modeling treatment-effect estimates from increasingly fine covariate partitions as Gaussian observations with bounded signal-to-noise ratios. We establish that, in general, the optimal treatment policy is not learnable in this setting. Even learning the optimal policy value suffers from impractically slow rates. However, when treatment effects vary smoothly, we derive minimax-adaptive policies based on linear smoothers that achieve vanishing welfare regret. We demonstrate the practical value of our framework by applying it to large-scale real-world experiments at Netflix, showing that personalized linear-smoothing policies can dominate unpersonalized policies even in this challenging empirical setting.
| Subjects: | Methodology (stat.ME); Machine Learning (cs.LG); Statistics Theory (math.ST) |
| Cite as: | arXiv:2610.10167 [stat.ME] |
| (or arXiv:2610.10167v1 [stat.ME] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10167 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Benedikt Koch [view email]
[v1]
Wed, 7 Oct 2026 14:41:58 UTC (248 KB)
来源:arXiv:cs.LG · arxiv.org