跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Wang Wei, Harry Yang, Tiankai Yang, Samyadeep Basu, Hongjie Chen, Andy Zhao, Franck Dernoncourt, Ryan A. Rossi, Hoda Eldardiry·· 5 小时前AI 评分34

FlexRouter:为灵活 LLM 路由学习互补模型集合

FlexRouter: Learning Complementary Model Sets for Flexible LLM Routing

AI 导读

FlexRouter 是一个显式建模模型互补性的 LLM 路由框架,以「答案覆盖」为目标,最大化所选模型中至少一个给出正确答案的概率。它用 Determinantal Point Processes(DPPs)建模路由策略,并通过边缘化失败集的训练目标直接优化覆盖,推理时用基于边缘对数行列式增益的贪心策略自适应决定子集大小。

正文

View PDF HTML (experimental)

Abstract:Existing Large Language Model (LLM) routing methods score LLMs independently to select top-$k$ models. However, this ignores model correlations and enforces a rigid computational budget. Consequently, routers often select redundant models that share failure modes, limiting the overall probability of success. To address this, we propose FlexRouter, a routing framework that explicitly models model complementarity. FlexRouter optimizes for \textit{answer coverage}, maximizing the probability that at least one selected model yields a correct response. This objective aligns with practical inference pipelines where multiple candidate outputs are generated and a downstream verifier or user selects the final one. We formulate routing as a coverage-oriented subset selection problem and model the routing policy using Determinantal Point Processes (DPPs), which naturally capture both model competence and redundancy. To directly optimize coverage without requiring a ground-truth target subset, we introduce a training objective based on marginalizing over failure sets. During inference, we employ a greedy strategy based on marginal log-determinant gains, enabling the router to adaptively determine subset sizes without a predefined budget. Extensive experiments on the large-scale RouterEval benchmark demonstrate that our proposed FlexRouter achieves higher coverage with lower redundancy across both in-domain and out-of-domain tasks than strong baselines while maintaining flexible inference cost.
Comments: 21 pages, 4 figures, accepted at COLM 2026
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as: arXiv:2609.38585 [cs.LG]
  (or arXiv:2609.38585v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.38585

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Wang Wei [view email]
[v1] Tue, 29 Sep 2026 21:44:05 UTC (315 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org