跳到正文
arXiv:cs.LG· Ian Osband·· 3 小时前AI 评分36

规划学习:horizon loss 如何改进交叉熵与精确策略梯度

Planning to Learn

AI 导读

研究提出 horizon loss,将交叉熵视为"耐心的准确率",把总误差按剩余学习量截断,从而在训练后期从交叉熵平滑过渡到精确策略梯度。在 MNIST 及 ImageNet 上使用 ResNet-50、ResNet-101 和 ViT-S/16,该方法在固定学习率下提升 top-1 准确率,且标签噪声越大增益越明显。

正文

View PDF HTML (experimental)

Abstract:Policy-gradient methods are central to modern reinforcement learning, including LLM post-training. When they struggle, the usual suspects are exploration, credit assignment and action-sampling noise. Classification has none of them. A classifier is a policy whose expected reward, its \emph{expected accuracy}, is the probability it assigns to the correct label, and because that label is known, the policy gradient is exact and smooth. Yet exact policy gradient loses to cross-entropy, even on expected accuracy. The exact gradient is myopic: it values an update only by what it buys now, but each update also sets where the next one starts, so an update's value depends on how much learning remains. Viewed this way, cross-entropy is patient accuracy, the total error an example would pay if its log-odds rose at unit speed forever, while exact policy gradient is the zero-horizon limit. Truncating this total at the learning that remains yields the horizon loss, a one-line change that moves from cross-entropy toward exact policy gradient as training runs out. In a simple allocation model, it provably escapes the trap that catches each endpoint. On MNIST and on ImageNet with ResNet-50, ResNet-101 and ViT-S/16, the horizon loss improves top-1 accuracy over cross-entropy at a flat learning rate, and the gain grows with label noise.
Subjects: Machine Learning (cs.LG); Optimization and Control (math.OC); Machine Learning (stat.ML)
Cite as: arXiv:2610.03667 [cs.LG]
  (or arXiv:2610.03667v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.03667

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Ian Osband [view email]
[v1] Fri, 2 Oct 2026 17:38:28 UTC (2,631 KB)

来源:arXiv:cs.LG · arxiv.org