跳到正文
arXiv:cs.AI· Ruoyu Qin, Zheming Li, Weiran He, Mingxing Zhang, Yongwei Wu, Weimin Zheng, Xinran Xu·· 3 小时前

Mooncake:面向 LLM 服务的 KVCache 中心化分离架构

Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

AI 导读

Moonshot AI 为 Kimi 打造的 LLM 服务推理平台 Mooncake 采用以 KVCache 为中心的分离式架构,将预填充与解码集群分离,并利用 GPU 集群中闲置的 CPU、DRAM 和 SSD 资源构建分离式 KVCache 缓存。

正文

View PDF HTML (experimental)

Abstract:Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI. It features a KVCache-centric disaggregated architecture that separates the prefill and decoding clusters. It also leverages the underutilized CPU, DRAM, and SSD resources of the GPU cluster to implement a disaggregated cache of KVCache. The core of Mooncake is its KVCache-centric scheduler, which balances maximizing overall effective throughput while meeting latency-related Service Level Objectives (SLOs). Unlike traditional studies that assume all requests will be processed, Mooncake faces challenges due to highly overloaded scenarios. To mitigate these, we developed a prediction-based early rejection policy. Experiments show that Mooncake excels in long-context scenarios. Compared to the baseline method, Mooncake can achieve up to a 525% increase in throughput in certain simulated scenarios while adhering to SLOs. Under real workloads, Mooncake's innovative architecture enables Kimi to handle 75% more requests.
Comments: 23 pages, 13 figures
Subjects: Distributed, Parallel, and Cluster Computing (cs.DC); Artificial Intelligence (cs.AI); Hardware Architecture (cs.AR)
Cite as: arXiv:2407.00079 [cs.DC]
  (or arXiv:2407.00079v5 [cs.DC] for this version)
  https://doi.org/10.48550/arXiv.2407.00079

arXiv-issued DOI via DataCite

Submission history

From: Ruoyu Qin [view email]
[v1] Mon, 24 Jun 2024 02:05:32 UTC (264 KB)
[v2] Tue, 2 Jul 2024 02:49:35 UTC (264 KB)
[v3] Tue, 9 Jul 2024 04:03:10 UTC (280 KB)
[v4] Wed, 3 Sep 2025 14:56:29 UTC (245 KB)
[v5] Thu, 8 Oct 2026 10:55:46 UTC (238 KB)

来源:arXiv:cs.AI · arxiv.org