ACE: Adaptive Calibration-Free Expert Skipping for MoE-based LLMs
Quick summary
arXiv:2609.05228v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) architectures provide an efficient paradigm for scaling large language models (LLMs), yet fixed top-k routing activates the same number of expert slots for every token, causing substantial redundant computation. Existing expert-skipping methods often rely on router confidence, calibration data, or additional training, and therefore cannot reliably estimate the actual contribution of routed experts. To this end, we propose ACE, a training-free, calibration-free, and checkpoint-preserving framework for token-adaptive expert
Key takeaways
- arXiv:2609.05228v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) architectures provide an efficient paradigm for scaling large language models (LLMs), yet fixed top-k routing activates the same number of expert slots for every token, causing substantial redundant computation.
- Existing expert-skipping methods often rely on router confidence, calibration data, or additional training, and therefore cannot reliably estimate the actual contribution of routed experts.
- To this end, we propose ACE, a training-free, calibration-free, and checkpoint-preserving framework for token-adaptive expert
Why it matters
The importance of “ACE: Adaptive Calibration-Free Expert Skipping for MoE-based LLMs” will be measured by what changes in practice. User behavior, access conditions, verifiable performance and responsible-use outcomes are the signals worth following.

Member comments