arXiv:cs.CL· Yibei Guo, Rui Liu·· 3 小时前AI 评分30
输入盲对照实验显示:多项选择题评估中层程序存在大量 oracle 提升空间
Input-Blind Controls Produce Substantial Oracle Headroom for Layer Programs in Multiple-Choice Evaluation
AI 导读
研究用 32 个跳层与重复程序、2 个模型和 4,413 道多选题,检验自适应计算中 oracle 选择带来的增益能否归因于所选层计算本身。在共享选项顺序下,输入盲扰动对照在 Qwen3-4B-Base 和 Llama-3.1-8B 上分别产生 10.2-11.8 与 15.6-19.4 个百分点的提升空间,超过真实程序的 9.0 和 10.1。
正文
Abstract:Adaptive computation aims to improve language-model inference by tailoring execution to each input. For layer programs, oracle evaluations use known answers to estimate the potential gain from this flexibility, before a practical selector is available. However, a gain from selection does not by itself explain why the chosen programs help. This study examines this distinction using 32 layer-skipping and repetition programs on two models and 4,413 multiple-choice items. The analysis compares their gains over a fixed action selected without the evaluation prompt with those of input-blind perturbations at the same sites, re-evaluating selections on another prompt. With shared option order, the controls give 10.2-11.8 and 15.6-19.4 percentage points of headroom on Qwen3-4B-Base and Llama-3.1-8B, exceeding the real programs' 9.0 and 10.1 in all three random-direction draws per model. They match answer-change rate only, and the ordering depends on the menu: in post hoc comparisons, real programs lead on Llama's repeat-only menu in every draw. A smaller KL-calibrated comparison, including an input-dependent control, favours real programs in point estimate, with inconclusive corrected tests. Fixed letter offsets produce headroom of similar scale. Rotating options sharply reduces both families' headroom, while leaving positive real-minus-control differences of 1.4-2.3 and 3.7-4.5 points; their magnitudes and statistical support depend on further adjustments and the reference. A supplementary generated-answer test finds that search-selected programs keep a 26.0-point advantage over programs selected for other problems after rewording, without a placebo comparison. These results show that substantial headroom can persist across prompts with shared option order without establishing a benefit specific to the selected layer computation; neither ordering against these controls identifies that benefit.
| Subjects: | Machine Learning (cs.LG); Computation and Language (cs.CL); Software Engineering (cs.SE) |
| Cite as: | arXiv:2610.10368 [cs.LG] |
| (or arXiv:2610.10368v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10368 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Rui Liu [view email]
[v1]
Wed, 7 Oct 2026 16:35:40 UTC (593 KB)
来源:arXiv:cs.CL · arxiv.org