Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Quick summary
arXiv:2407.00079v5 Announce Type: replace-cross Abstract: Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI. It features a KVCache-centric disaggregated architecture that separates the prefill and decoding clusters. It also leverages the underutilized CPU, DRAM, and SSD resources of the GPU cluster to implement a disaggregated cache of KVCache. The core of Mooncake is its KVCache-centric scheduler, which balances maximizing overall effective throughput while meeting latency-related Service Level Objectives (SLOs). Unlike traditional studies that assume al
Key takeaways
- arXiv:2407.00079v5 Announce Type: replace-cross Abstract: Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
- It features a KVCache-centric disaggregated architecture that separates the prefill and decoding clusters.
- It also leverages the underutilized CPU, DRAM, and SSD resources of the GPU cluster to implement a disaggregated cache of KVCache.
Why it matters
AI progress is not only a software story. Chips, data centers and energy decisions help determine which models can operate economically and what end users ultimately pay.

Member comments