arXiv Artificial Intelligence

Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Quick summary

arXiv:2407.00079v5 Announce Type: replace-cross Abstract: Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI. It features a KVCache-centric disaggregated architecture that separates the prefill and decoding clusters. It also leverages the underutilized CPU, DRAM, and SSD resources of the GPU cluster to implement a disaggregated cache of KVCache. The core of Mooncake is its KVCache-centric scheduler, which balances maximizing overall effective throughput while meeting latency-related Service Level Objectives (SLOs). Unlike traditional studies that assume al

Key takeaways

  • arXiv:2407.00079v5 Announce Type: replace-cross Abstract: Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
  • It features a KVCache-centric disaggregated architecture that separates the prefill and decoding clusters.
  • It also leverages the underutilized CPU, DRAM, and SSD resources of the GPU cluster to implement a disaggregated cache of KVCache.

Why it matters

AI progress is not only a software story. Chips, data centers and energy decisions help determine which models can operate economically and what end users ultimately pay.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗