arXiv Artificial Intelligence

APEX: Adaptive Expert Prefetching for Memory-Efficient Edge MoE Inference

APEX: Adaptive Expert Prefetching for Memory-Efficient Edge MoE Inference

Quick summary

arXiv:2608.11688v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) models are attractive for edge deployment because they provide high model capacity while activating only a small subset of parameters per token, improving compute efficiency. However, MoE inference at the edge is fundamentally limited by memory. Expert parameters are large and often reside in off-chip memory due to capacity, cost, and power constraints, putting expert loading to the critical path. We present APEX: Adaptive Expert Prefetching, a predictive resource management framework that overlaps expert loading with u

Key takeaways

  • arXiv:2608.11688v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) models are attractive for edge deployment because they provide high model capacity while activating only a small subset of parameters per token, improving compute efficiency.
  • However, MoE inference at the edge is fundamentally limited by memory.
  • Expert parameters are large and often reside in off-chip memory due to capacity, cost, and power constraints, putting expert loading to the critical path.

Why it matters

“APEX: Adaptive Expert Prefetching for Memory-Efficient Edge MoE Inference” exposes the compute, energy and supply-chain layer behind model competition. Capacity shifts can influence model costs, service availability and the ability of smaller companies to compete.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗