arXiv Artificial Intelligence

ACE: Adaptive Calibration-Free Expert Skipping for MoE-based LLMs

ACE: Adaptive Calibration-Free Expert Skipping for MoE-based LLMs

Quick summary

arXiv:2609.05228v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) architectures provide an efficient paradigm for scaling large language models (LLMs), yet fixed top-k routing activates the same number of expert slots for every token, causing substantial redundant computation. Existing expert-skipping methods often rely on router confidence, calibration data, or additional training, and therefore cannot reliably estimate the actual contribution of routed experts. To this end, we propose ACE, a training-free, calibration-free, and checkpoint-preserving framework for token-adaptive expert

Key takeaways

  • arXiv:2609.05228v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) architectures provide an efficient paradigm for scaling large language models (LLMs), yet fixed top-k routing activates the same number of expert slots for every token, causing substantial redundant computation.
  • Existing expert-skipping methods often rely on router confidence, calibration data, or additional training, and therefore cannot reliably estimate the actual contribution of routed experts.
  • To this end, we propose ACE, a training-free, calibration-free, and checkpoint-preserving framework for token-adaptive expert

Why it matters

The importance of “ACE: Adaptive Calibration-Free Expert Skipping for MoE-based LLMs” will be measured by what changes in practice. User behavior, access conditions, verifiable performance and responsible-use outcomes are the signals worth following.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗