arXiv:cs.LG· Nicol\`o Felicioni, Michael Benigni, Maurizio Ferrari Dacrema, Paolo Cremonesi·· 4 小时前AI 评分28
VOCEM:基于联合效应建模的方差最优离策略评估
Variance-Optimal Off-Policy Evaluation with Conjunct Effect Modeling
AI 导读
针对离策略评估中动作级重要性权重方差过大的问题,研究者提出 VOCEM 估计器,在 DR 与 OffCEM 之间插值并以闭式解选取最小化方差的系数,理论上方差不超过两者端点。在受控合成实验和两个大规模动作基准上,VOCEM 在全部 23 种评估条件下均优于 DR 与 OffCEM,稳定性与鲁棒性更好。
正文
Abstract:Off-policy evaluation (OPE) for contextual bandit policies becomes challenging when action-level importance weighting incurs excessive variance. Doubly robust (DR) estimation remains unbiased under common support but retains these high-variance action-level weights. A prior estimator, Off-policy evaluation with Conjunct Effect Model (OffCEM), replaces them with more stable cluster-level weights, at the cost of relying on local correctness of the reward model. In this paper, we show that, under the assumptions required by DR and OffCEM, there exists an unbiased family of estimators that interpolates between OffCEM and DR. Building on this result, we propose the Variance Optimal-CEM (VOCEM) estimator, which selects the interpolation coefficient to minimize variance. We derive the population-optimal coefficient in closed form and show that the resulting estimator has variance no larger than either endpoint, OffCEM or DR. Experiments in controlled synthetic settings and on two large-action benchmarks show that VOCEM improves upon both endpoints in all 23 evaluated conditions, exhibiting greater stability and empirical robustness.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.08677 [cs.LG] |
| (or arXiv:2610.08677v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08677 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Michael Benigni [view email]
[v1]
Tue, 6 Oct 2026 16:58:26 UTC (1,281 KB)
来源:arXiv:cs.LG · arxiv.org