arXiv Artificial Intelligence

UniMoMo: Expert Merging-Based MoE Acceleration for Large Recommendation Models

UniMoMo: Expert Merging-Based MoE Acceleration for Large Recommendation Models

Quick summary

arXiv:2608.08627v1 Announce Type: new Abstract: Sparse mixture-of-experts (MoE) layers expand recommendation capacity through conditional computation, yet a trained checkpoint still stores and routes over its full expert bank. We study a deployment problem: convert that checkpoint to a smaller standard MoE under an explicit expert budget, without adding a compression-specific online module. To address this, we introduce UniMoMo, a post-training compression framework formulated as a constrained graph coarsening problem. Rather than relying on parameter distance, UniMoMo groups experts based on

Key takeaways

  • arXiv:2608.08627v1 Announce Type: new Abstract: Sparse mixture-of-experts (MoE) layers expand recommendation capacity through conditional computation, yet a trained checkpoint still stores and routes over its full expert bank.
  • We study a deployment problem: convert that checkpoint to a smaller standard MoE under an explicit expert budget, without adding a compression-specific online module.
  • To address this, we introduce UniMoMo, a post-training compression framework formulated as a constrained graph coarsening problem.

Why it matters

“UniMoMo: Expert Merging-Based MoE Acceleration for Large Recommendation Models” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗