跳到正文
arXiv:cs.LG· Jungsoo Park, Hyungjoo Chae, Ethan Mendes, Jay DeYoung, Varsha Kishore, Wei Xu, Alan Ritter·· 7 小时前AI 评分35

面向 LLM 回归的预测分布强化学习:DAR 方法

Reinforcement Learning over Predictive Distributions for LLM Regression

AI 导读

研究者提出 Distribution-Aware Reward(DAR),一种 on-policy 强化学习目标,通过评估同一输入多次预测形成的经验预测分布,并按每个预测对整体分布质量的 leave-one-out 贡献分配奖励。在合成插值/外推任务和涉及代码、分子数据的两个真实科学回归任务上,DAR 相比监督微调和逐点强化学习,产生了校准更好的不确定性估计,同时持续降低预测误差并提升排序质量。

正文

View PDF HTML (experimental)

Abstract:Large language models (LLMs) have emerged as flexible regressors capable of predicting real-valued quantities from heterogeneous inputs. Yet most LLM regression objectives optimize predictions independently, often yielding poor calibration. We introduce Distribution-Aware Reward (DAR), an on-policy reinforcement learning objective that instead jointly evaluates the empirical predictive distribution formed by multiple predictions for the same input. To translate this distribution-level objective into rollout-level rewards, we assign each prediction credit based on its leave-one-out contribution to the quality of the overall predictive distribution. This encourages predictions that are well-centered and appropriately dispersed around the target. We evaluate on three regression settings: a synthetic task probing interpolation and extrapolation, and two real-world scientific tasks involving code and molecular data. Across tasks, DAR produces better-calibrated uncertainty estimates while consistently reducing prediction error and improving ranking quality over supervised fine-tuning and pointwise reinforcement learning. Together, these results highlight the benefits of distribution-aware training for LLM regression.
Comments: 27 pages, 7 figures
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as: arXiv:2605.20740 [cs.LG]
  (or arXiv:2605.20740v2 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2605.20740

arXiv-issued DOI via DataCite

Submission history

From: Jungsoo Park [view email]
[v1] Wed, 20 May 2026 05:43:40 UTC (2,468 KB)
[v2] Tue, 6 Oct 2026 17:48:09 UTC (3,495 KB)

来源:arXiv:cs.LG · arxiv.org