arXiv:cs.LG· Younwoo Choi, Leo Feng, Vincent Liu, Haanvid Lee·· 3 小时前
PFN-OPE:面向 LLM 的摊销式离策略评估
Amortized Off-Policy Evaluation for LLMs
AI 导读
研究者提出 PFN-OPE,一种先验数据拟合网络,将离策略评估(OPE)摊销到一类上下文赌博机任务上,单次前向传播即可由日志数据和每个提示的一个目标回复映射出价值估计,无需按任务重新拟合。在 HelpSteer2 与 UltraFeedback 上、针对 Qwen、Llama 与 Gemma 策略,奖励偏移场景下其误差比最佳基线低 2.0 至 9.3 倍。
正文
Abstract:Accurate evaluation is central to selecting which LLM to deploy, yet testing a candidate on live traffic exposes real users to an unvetted model. Teams therefore evaluate candidates offline, on data produced by already-deployed models. This is off-policy evaluation (OPE), and it faces two distribution shifts: as a model is updated in post-training, its responses diverge from the logged ones (policy shift), and the reward definition under which it is judged changes with business requirements (reward shift). Classical OPE methods are ill-suited to this continual-deployment setting because they are defined per task and require fitting from scratch on every new logged dataset or reward definition. To address this, we propose PFN-OPE, a prior-data fitted network that amortizes OPE across a distribution of contextual-bandit tasks. We pretrain it once on tasks constructed from a pool of LLM responses scored by several reward functions, in which both shifts occur. At test-time it maps a logged dataset and one sampled target response per prompt to a value estimate in a single forward pass, with no per-task fitting. On HelpSteer2 and UltraFeedback with Qwen, Llama, and Gemma policies, PFN-OPE achieves 2.0 to 9.3 times lower error than the best baselines across all tested configurations in the reward-shifted settings.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.10848 [cs.LG] |
| (or arXiv:2610.10848v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10848 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Younwoo Choi Mr. [view email]
[v1]
Wed, 7 Oct 2026 19:51:59 UTC (463 KB)
来源:arXiv:cs.LG · arxiv.org