arXiv:cs.LG· Indranil Halder, Cengiz Pehlevan·· 5 小时前AI 评分48
解析 LLM-as-a-Judge:面向推理时扩展的可解析模型
Demystifying LLM-as-a-Judge: Analytically Tractable Model for Inference-Time Scaling
AI 导读
研究者提出一个可解析的推理时扩展模型——带奖励加权采样的贝叶斯线性回归,用于刻画 LLM-as-a-Judge 场景。当奖励与教师模型差距不大时,泛化误差随推理采样数 k 单调下降;但奖励严重错配会带来有限的最优 k,超过后更多采样反而增大误差。在"best-of-k"极限下,以教师为奖励时泛化误差按 Θ(1/k²) 衰减,且任务难度上升会削弱推理时算力的优势。
正文
Abstract:Recent developments in large language models have shown advantages in reallocating a notable share of computational resource from training time to inference time. However, the principles behind inference time scaling are not well understood. In this paper, we introduce an analytically tractable model of inference-time scaling: Bayesian linear regression with a reward-weighted sampler, where the reward is determined from a linear model, modeling LLM-as-a-judge scenario. We study this problem in the high-dimensional regime, where the deterministic equivalents dictate a closed-form expression for the posterior predictive mean and variance. We analyze the generalization error when training data are sampled from a teacher model. We draw $k$ inference-time samples and select via softmax at a temperature applied to a quadratic reward. When the reward is not too different from the teacher, the generalization error decreases monotonically with increasing inference time samples $k$. However, the specific reward that optimizes inference-time selection generally differs from the teacher. In contrast, substantial reward misspecification induces a finite optimal $k$ beyond which more sampling can increase the generalization error. For fixed $k$, there exists an optimal sampling temperature. We experimentally verify these facts in large language model inference with an additional large language model as a judge. In the "best-of-$k$" limit with the teacher as reward, we theoretically show that the generalization error decays as $\Theta(1/k^2)$ and determine the leading coefficient via extreme value theory. These formulas delineate domains where scaling inference-time computation is provably preferable to collecting more data. Finally, we demonstrate that when task difficulty increases, the previously mentioned advantage of inference-time compute degrades.
| Comments: | Published at International Conference on Machine Learning 2026 |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2512.19905 [cs.LG] |
| (or arXiv:2512.19905v3 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2512.19905 arXiv-issued DOI via DataCite |
Submission history
From: Indranil Halder [view email]
[v1]
Mon, 22 Dec 2025 22:13:06 UTC (6,657 KB)
[v2]
Wed, 11 Feb 2026 21:21:08 UTC (8,569 KB)
[v3]
Fri, 2 Oct 2026 16:16:57 UTC (8,588 KB)
来源:arXiv:cs.LG · arxiv.org