arXiv:cs.LG(机器学习,全量分类)· Wang Wei, Harry Yang, Tiankai Yang, Samyadeep Basu, Hongjie Chen, Andy Zhao, Franck Dernoncourt, Ryan A. Rossi, Hoda Eldardiry·· 5 小时前AI 评分34
FlexRouter:为灵活 LLM 路由学习互补模型集合
FlexRouter: Learning Complementary Model Sets for Flexible LLM Routing
AI 导读
FlexRouter 是一个显式建模模型互补性的 LLM 路由框架,以「答案覆盖」为目标,最大化所选模型中至少一个给出正确答案的概率。它用 Determinantal Point Processes(DPPs)建模路由策略,并通过边缘化失败集的训练目标直接优化覆盖,推理时用基于边缘对数行列式增益的贪心策略自适应决定子集大小。
正文
Abstract:Existing Large Language Model (LLM) routing methods score LLMs independently to select top-$k$ models. However, this ignores model correlations and enforces a rigid computational budget. Consequently, routers often select redundant models that share failure modes, limiting the overall probability of success. To address this, we propose FlexRouter, a routing framework that explicitly models model complementarity. FlexRouter optimizes for \textit{answer coverage}, maximizing the probability that at least one selected model yields a correct response. This objective aligns with practical inference pipelines where multiple candidate outputs are generated and a downstream verifier or user selects the final one. We formulate routing as a coverage-oriented subset selection problem and model the routing policy using Determinantal Point Processes (DPPs), which naturally capture both model competence and redundancy. To directly optimize coverage without requiring a ground-truth target subset, we introduce a training objective based on marginalizing over failure sets. During inference, we employ a greedy strategy based on marginal log-determinant gains, enabling the router to adaptively determine subset sizes without a predefined budget. Extensive experiments on the large-scale RouterEval benchmark demonstrate that our proposed FlexRouter achieves higher coverage with lower redundancy across both in-domain and out-of-domain tasks than strong baselines while maintaining flexible inference cost.
| Comments: | 21 pages, 4 figures, accepted at COLM 2026 |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2609.38585 [cs.LG] |
| (or arXiv:2609.38585v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2609.38585 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Wang Wei [view email]
[v1]
Tue, 29 Sep 2026 21:44:05 UTC (315 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org