arXiv Artificial Intelligence

RouterInterp: Understanding Superposed Specialisation in Mixture of Experts Routing

RouterInterp: Understanding Superposed Specialisation in Mixture of Experts Routing

Quick summary

arXiv:2610.11775v1 Announce Type: new Abstract: Sparse Mixture of Experts (MoE) models scale more efficiently than dense models by routing tokens to modular expert networks that are only active for processing a fraction of tokens. A leading hypothesis for the performance of MoE models is that each expert specialises in a single, coherent domain. However, interpretability efforts that assume this hypothesis have generally been unsuccessful. We propose and present evidence for an alternative account that we call the Superposed Specialisation Hypothesis (SSH): experts specialise in a disjoint uni

Key takeaways

  • arXiv:2610.11775v1 Announce Type: new Abstract: Sparse Mixture of Experts (MoE) models scale more efficiently than dense models by routing tokens to modular expert networks that are only active for processing a fraction of tokens.
  • A leading hypothesis for the performance of MoE models is that each expert specialises in a single, coherent domain.
  • However, interpretability efforts that assume this hypothesis have generally been unsuccessful.

Why it matters

The importance of “RouterInterp: Understanding Superposed Specialisation in Mixture of Experts Routing” will be measured by what changes in practice. User behavior, access conditions, verifiable performance and responsible-use outcomes are the signals worth following.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗