跳到正文
Rohan Paul· @rohanpaul_ai · X·· 2 小时前AI 评分43
AI 导读

清华提出 TokenRouter 服务系统,在 token 级做大小模型路由,吞吐量最高达现有方案的 64.15 倍。它让每个模型独享服务器并来回传递半成品答案,同时保留 KV cache、短暂缓存请求以凑更大批次。在 5 种路由方法上,吞吐量比更强的现有方案提升 2.01 至 64.15 倍。

正文

New Tsinghua paper builds TokenRouter, a serving system that runs per-token small-and-large model routing at upto 64.15X the throughput of existing setups.

Current popular serving frameworks (like vLLM and SGLang) run one model per request, so when two models share an answer, every step waits for the slower one.

TokenRouter gives each model its own server and lets them pass work back and forth. So, it hands a half-written answer between them while keeping the model’s memory of the text so far (the KV cache), and holds requests for a moment so each model works on bigger batches.

Across 5 routing methods, throughput rose 2.01 to 64.15 times over the stronger existing setup.

– arxiv. org/abs/2610.12242

Title: "TokenRouter: Efficient Serving System for Token-Level LLM Routing"

来源:Rohan Paul · x.com