跳到正文
arXiv:cs.AI· Sam Wang, Julia White, Sahibzada Allahyar, Dhruv Atreja, Urchade Zaratiana, Kelton Zhang·· 5 小时前AI 评分57

研究分析六个商业 LLM 路由器:多数表现不如随机选择两个模型

Dynamic LLM Routers are Often Misguided

AI 导读

arXiv 论文(arXiv:2610.02762)分析六个商业 LLM 路由器在 14 个设置、八类任务基准上的表现,发现没有一个在同等成本下胜过在两个精选模型间随机选择的路由器,部分落后超过 10 个百分点。

正文

View PDF HTML (experimental)

Abstract:Dynamic LLM routers promise to cut inference costs by sending each query to the cheapest model that can answer it correctly. We analyze six commercial routers across 14 settings on a diverse benchmark spanning eight task categories, finding that none of them outperforms a router that randomly selects between two well-chosen models at matched cost. Some underperform by more than 10 percentage points. We trace this gap to four patterns prevalent across routers: difficulty blindness, length reversal, semantic matching, and roster suboptimality. We show that the first three are what the standard objective rewards: cost-accuracy Pareto efficiency on realized costs favors escalating moderately hard queries over the hardest ones, shorter queries over longer ones, and routing by a query's source over its difficulty. We also argue that the two assumptions that would justify large rosters, model granularity and model specialization, do not hold empirically. We propose an alternative evaluation methodology that does not reward these patterns, and as a proof of concept, we design a simple two-model router that avoids all four. Nevertheless, its gain over random routing is limited, because a well-chosen roster leaves little to route.
Comments: 8 pages, 15 figures, under review at NAACL
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.02762 [cs.AI]
  (or arXiv:2610.02762v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.02762

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Sam Wang [view email]
[v1] Fri, 2 Oct 2026 03:42:28 UTC (565 KB)

来源:arXiv:cs.AI · arxiv.org