跳到正文
arXiv:cs.AI· Ramin Pishehvar, Andrea Morandi, Mahesh Viswanathan·· 6 小时前AI 评分45

如何评估 LLM 路由的升级信号:目标、对照与五种自欺方式

Evaluating Escalation Signals for LLM Routing: Targets, Controls, and Five Ways to Fool Yourself

AI 导读

研究测试语义熵作为 LLM 路由升级信号,在 GSM8K 上大小模型约 12 倍参数差距时以 AUROC 0.871 识别小模型错误,匹配成本下路由准确率比随机升级最高提升 9 个百分点。

正文

View PDF HTML (experimental)

Abstract:Deciding when to escalate a query from a small language model to a larger one requires a cheap signal that predicts, before the large model is called, whether escalating would help. Semantic entropy, originally developed to detect hallucinations, is a natural candidate: it measures how much a model's sampled answers disagree in meaning, and high disagreement often signals an unreliable answer. We test it across three benchmarks and two model families. On GSM8K, with a small/large pair about twelve times apart in size, semantic entropy reliably distinguishes the small model's mistakes (AUROC 0.871) and improves routed accuracy over random escalation by up to nine points at matched cost. An earlier strong-looking result on a synthetic benchmark proved misleading: a simple rule based only on question difficulty, with no model involved, matched semantic entropy almost exactly. This paper's main contribution is a set of checks that catch this before it is reported as real. We show that scoring a cheap, question-only difficulty estimate alongside any signal reveals whether the signal adds real information or just tracks how hard a question looks; that two reasonable definitions of "escalation worked" can produce very different results on the same data; that a benchmark can leave almost no room for any signal to beat simply always using the large model; and that the true cost of live sampling can make routing more expensive than calling the large model directly. For a cheaper alternative that reuses cached past outcomes, we show how to predict whether it will work on a new dataset -- confirmed by correctly forecasting a collapse from AUROC 0.908 to chance level (0.518) ahead of time. We offer these as a general checklist for evaluating escalation signals.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.07354 [cs.AI]
  (or arXiv:2610.07354v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.07354

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Ramin Pishehvar [view email]
[v1] Mon, 5 Oct 2026 20:22:34 UTC (240 KB)

来源:arXiv:cs.AI · arxiv.org