arXiv:cs.CL· Manh Nguyen, Sunil Gupta, Hung Le·· 3 小时前
UAB:面向自适应测试时推理的不确定性感知预算分配
Uncertainty-Aware Budget Allocation for Adaptive Test-Time Reasoning
AI 导读
UAB 是一种不确定性感知预算分配框架,通过初始采样答案的投票熵估计难度,再用边际贪心算法把剩余采样预算集中到答案分歧大的难题上。在五个 1.5B–27B 开源模型和五个推理基准上,UAB 平均准确率比均匀分配高 +2.3%,单项基准最高提升 +5.5%,低资源场景收益最大。该方法精确用完目标预算,无需辅助模型或额外 LLM 调用,代码已公开。
正文
Abstract:Sampling multiple responses improves language model reasoning, but uniform compute allocation is inefficient because easy questions are over-sampled while hard questions remain under-explored. We propose \textbf{Uncertainty-Aware Budget Allocation (UAB)}, a concave integer optimization framework that reallocates a fixed sampling budget using uncertainty estimated from the initial samples themselves. In Phase-1, every question receives a small fixed number of generations. Their answer disagreement, measured by vote entropy, provides a difficulty signal while these generations contribute to the final vote. In Phase-2, the remaining budget is allocated by a marginal-greedy algorithm that optimally solves a concave coverage-maximization surrogate, concentrating samples on questions whose initial answers disagree. Across five open-weight models (1.5B--27B parameters) and five reasoning benchmarks of varying difficulty, UAB improves average accuracy by $+2.3\%$ over uniform allocation, and achieves gains of up to $+5.5\%$ on individual benchmarks, with the largest gains in low-resource settings. Moreover, UAB meets the target budget exactly and requires no auxiliary model or additional LLM calls. Code is publicly available at this https URL.
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2605.26849 [cs.CL] |
| (or arXiv:2605.26849v2 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2605.26849 arXiv-issued DOI via DataCite |
Submission history
From: Manh Nguyen [view email]
[v1]
Tue, 26 May 2026 11:06:58 UTC (212 KB)
[v2]
Thu, 8 Oct 2026 10:48:35 UTC (180 KB)
来源:arXiv:cs.CL · arxiv.org