Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE
Quick summary
arXiv:2609.38140v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE), popularized by large language models, is a promising paradigm for scaling visual generative models. However, conventional token-wise MoE routes tokens independently within a homogeneous expert pool and regularizes expert usage toward uniformity, making it poorly matched to video data that is spatiotemporally redundant and semantically long-tailed. We show that existing visual MoEs fall into a uniformity trap: semantically under-organized routing, compounded by uniform expert-usage regularization, scatters coherent patc
Key takeaways
- arXiv:2609.38140v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE), popularized by large language models, is a promising paradigm for scaling visual generative models.
- However, conventional token-wise MoE routes tokens independently within a homogeneous expert pool and regularizes expert usage toward uniformity, making it poorly matched to video data that is spatiotemporally redundant and semantically long-tailed.
- We show that existing visual MoEs fall into a uniformity trap: semantically under-organized routing, compounded by uniform expert-usage regularization, scatters coherent patc
Why it matters
“Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Member comments