arXiv:cs.AI· Ruoyu Qin, Zheming Li, Weiran He, Mingxing Zhang, Yongwei Wu, Weimin Zheng, Xinran Xu·· 3 小时前
Mooncake:面向 LLM 服务的 KVCache 中心化分离架构
Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
AI 导读
Moonshot AI 为 Kimi 打造的 LLM 服务推理平台 Mooncake 采用以 KVCache 为中心的分离式架构,将预填充与解码集群分离,并利用 GPU 集群中闲置的 CPU、DRAM 和 SSD 资源构建分离式 KVCache 缓存。
正文
Abstract:Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI. It features a KVCache-centric disaggregated architecture that separates the prefill and decoding clusters. It also leverages the underutilized CPU, DRAM, and SSD resources of the GPU cluster to implement a disaggregated cache of KVCache. The core of Mooncake is its KVCache-centric scheduler, which balances maximizing overall effective throughput while meeting latency-related Service Level Objectives (SLOs). Unlike traditional studies that assume all requests will be processed, Mooncake faces challenges due to highly overloaded scenarios. To mitigate these, we developed a prediction-based early rejection policy. Experiments show that Mooncake excels in long-context scenarios. Compared to the baseline method, Mooncake can achieve up to a 525% increase in throughput in certain simulated scenarios while adhering to SLOs. Under real workloads, Mooncake's innovative architecture enables Kimi to handle 75% more requests.
| Comments: | 23 pages, 13 figures |
| Subjects: | Distributed, Parallel, and Cluster Computing (cs.DC); Artificial Intelligence (cs.AI); Hardware Architecture (cs.AR) |
| Cite as: | arXiv:2407.00079 [cs.DC] |
| (or arXiv:2407.00079v5 [cs.DC] for this version) | |
| https://doi.org/10.48550/arXiv.2407.00079 arXiv-issued DOI via DataCite |
Submission history
From: Ruoyu Qin [view email]
[v1]
Mon, 24 Jun 2024 02:05:32 UTC (264 KB)
[v2]
Tue, 2 Jul 2024 02:49:35 UTC (264 KB)
[v3]
Tue, 9 Jul 2024 04:03:10 UTC (280 KB)
[v4]
Wed, 3 Sep 2025 14:56:29 UTC (245 KB)
[v5]
Thu, 8 Oct 2026 10:55:46 UTC (238 KB)
来源:arXiv:cs.AI · arxiv.org