跳到正文
arXiv:cs.CL· Ran Li, Lei Chen·· 6 小时前AI 评分40

Prefill-Only 决策模型的读出稳定性:零标签预测与推理时算力分配

Readout Stability in Prefill-Only Decision Models:Zero-Label Prediction and Inference-Time Compute Allocation

AI 导读

受 Jev 模型启发的 prefill-only 决策模型每次调用比同规模生成式语言模型便宜一到两个数量级,其首次前向传播的缓存分布已决定候选菜单变动后的准确率。该估计器无需标签、无需二次前向传播,在七个模型家族、十个数据集和两类任务上误差在 4.2 分以内。在精选 5 候选菜单上,0.8B 模型在 CLINC150 达到 95.4%,高于 4B 模型在全部 150 标签上的 80.0%。

正文

View PDF HTML (experimental)

Abstract:Prefill-only decision models inspired by the Jev model score every candidate in a menu during a single forward pass and never decode, which makes one call one to two orders of magnitude cheaper than a same-scale generative language model. We show that this read-out structure comes with a testable property. When an intervention changes only the candidate menu and leaves the input text fixed, the post-intervention accuracy is already determined by the cached first-pass distribution. The estimator restricts the pass-1 probabilities to the menu, renormalizes, and reads off the argmax; it uses no labels and no second forward pass. Across seven model families, ten datasets and two task types, menu-only interventions are predicted to within 4.2 points, and for one family the prediction is exact. A probability-level variant of the same estimator errs by 21.0 points, so the property lives in the ranking rather than in the probabilities and is not recovered by calibration. Same-scale generative language models do not share the property. On those models the same estimator errs by 1.6 to 15.8 points and degrades as the model grows. The property turns inference-time compute into a decision that can be made before deployment. Uniform extra passes buy calibration but almost no accuracy; at matched cost a confidence cascade outperforms every scheme that re-asks the same model, and curating the menu beats enlarging the model, with a 0.8B model on a curated 5-candidate menu reaching 95.4% on CLINC150 against 80.0% for a 4B model on the full 150-label this http URL and data are available at this https URL.
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2610.07716 [cs.CL]
  (or arXiv:2610.07716v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.07716

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Ran Li [view email]
[v1] Tue, 6 Oct 2026 04:10:16 UTC (1,483 KB)

来源:arXiv:cs.CL · arxiv.org