arXiv:cs.AI· Biao Xiang, Ali Eshragh, Yuexing Li, Kai Wang·· 3 小时前
面向离策略评估的预算化多源反事实标注
Budgeted Multi-Source Counterfactual Annotation for Off-Policy Evaluation
AI 导读
研究针对上下文赌博机离策略评估,在给定各标注源成本与误差特征下,将上下文-动作对与标注源的分配建模为整数规划问题,以最小化依赖标注方案的估计量方差分量,并提出带动态规划子程序的优化算法单调改进目标。在合成临床与 LLM 标注教育赌博机实验中,该方法相比无标注分别将固定方案 MSE 降低 20.58% 和 10.77%,论文已被 NeurIPS 2026 接收。
正文
Abstract:Off-policy evaluation (OPE) estimates the value of a target policy from logged data, but limited behavior-policy coverage can force high-variance reweighting or reward-model extrapolation. Counterfactual annotations can add evidence about unobserved actions, yet practical sources, including domain experts and large language models (LLMs), may be costly, biased, or noisy. We study budgeted acquisition of such annotations for contextual-bandit OPE. Given source-specific costs and error profiles, we formulate an integer allocation problem over context-action pairs and annotation sources to minimize the component of estimator variance that depends on the annotation plan. We characterize when annotations are valuable through a first-annotation threshold and local annotation-value regimes. For the coupled multi-source problem, we develop a majorization-minimization algorithm with dynamic-programming subroutines that monotonically improves the objective. Experiments in synthetic clinical and LLM-annotated education bandits show that our allocation method reduces fixed-profile mean squared error (MSE) by 20.58% and 10.77%, respectively, relative to no annotation.
| Comments: | Accepted to NeurIPS 2026. Code available at this https URL |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Optimization and Control (math.OC) |
| Cite as: | arXiv:2610.10974 [cs.LG] |
| (or arXiv:2610.10974v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10974 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Biao Xiang [view email]
[v1]
Wed, 7 Oct 2026 22:55:15 UTC (662 KB)
来源:arXiv:cs.AI · arxiv.org