arXiv Artificial Intelligence

EMA: Elastic and Performance Transparent Memory Across GPUs

EMA: Elastic and Performance Transparent Memory Across GPUs

Quick summary

arXiv:2609.27040v1 Announce Type: cross Abstract: Multi-GPU servers have become the standard building block of modern data centers, providing aggregated capacity through high-bandwidth interconnects. At the same time, workloads such as LLM inference exhibit highly dynamic memory demands, which can cause one GPU to exhaust its local memory while others remain underutilized. This mismatch motivates a model of elastic resource sharing across GPUs. We present EMA, a memory sharing system that allows GPUs within a server to borrow and reclaim memory from each other, forming an elastic pool of capac

Key takeaways

  • arXiv:2609.27040v1 Announce Type: cross Abstract: Multi-GPU servers have become the standard building block of modern data centers, providing aggregated capacity through high-bandwidth interconnects.
  • At the same time, workloads such as LLM inference exhibit highly dynamic memory demands, which can cause one GPU to exhaust its local memory while others remain underutilized.
  • This mismatch motivates a model of elastic resource sharing across GPUs.

Why it matters

AI progress is not only a software story. Chips, data centers and energy decisions help determine which models can operate economically and what end users ultimately pay.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗