跳到正文
HuggingFace Daily Papers·· 2 天前

TokenRouter:面向 Token 级 LLM 路由的高效服务系统

TokenRouter: Efficient Serving System for Token-Level LLM Routing

AI 导读

TokenRouter 是一个面向 Token 级 LLM 路由的高效服务系统,遵循「请求中心编程、模型中心执行」原则,为每个 LLM 启动子服务器并异步调度请求,其延迟批处理调度器超参数由吞吐数学模型推导。在多种路由算法、负载和模型组合下,TokenRouter 的解码吞吐量比现有系统高 2.01-64.15 倍。该工作已被 NeurIPS 2026 接收,代码已开源。

正文

View PDF HTML (experimental)

Abstract:Large language model (LLM) routing distributes inference work across different models, advancing the cost-quality Pareto frontier of LLM serving. While coarse-grained routing at the session or query level has been widely adopted in production systems, recent algorithmic work shows that fine-grained token-level routing can yield substantial efficiency and quality gains. However, efficiently serving token-level routed inference poses significant challenges to existing systems. Built on single-LLM assumptions, current systems suffer from severe step desynchronization and frequent batch admission delays under token-level routing, and they also impose high implementation complexity on developers. To address these challenges, we design TokenRouter, an efficient and developer-friendly serving system for token-level routed LLM inference. TokenRouter follows the principle of request-centric programming, model-centric execution: developers describe routing logic from the perspective of a single request, while the runtime launches a subserver for each LLM and dispatches requests asynchronously. Each subserver employs a delayed-batching scheduler, whose optimal hyperparameters are derived from a mathematical throughput model of the system. Across diverse routing algorithms, workloads, and model pairs, TokenRouter achieves 2.01-64.15x higher decoding throughput than existing systems, substantially advancing the serving efficiency of token-level LLM routing. Our code is available at this https URL.
Comments: Accepted by NeurIPS 2026
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2610.12242 [cs.CL]
  (or arXiv:2610.12242v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.12242

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Tengxuan Liu [view email]
[v1] Thu, 8 Oct 2026 16:21:50 UTC (329 KB)

来源:HuggingFace Daily Papers · arxiv.org