arXiv:2610.00558v1 Announce Type: cross Abstract: While Mixture-of-Experts (MoE) models effectively scale model capacity through sparse activation, their deployment is often bottlenecked by prohibitive memory requirements. Extracting a compact subset of experts presents a promising solution. However, existing expert selection heuristics predominantly rely on Top-k ranking, which isolates the eval
DS-MoE is a theoretically grounded framework for Mixture-of-Experts (MoE) model deployment that addresses memory bottlenecks by selecting compact expert subsets via difference-of-submodular (DS) optimization. Unlike standard Top-k heuristics that ignore inter-expert dependencies, DS-MoE leverages the MoE Hessian to quantify and balance functional redundancy (overlapping representations) and synergy (complementary error cancellation).
The method employs a tailored majorization-minimization (MM) algorithm with provable monotonic convergence to efficiently identify optimal expert combinations. This approach significantly improves training efficiency, such as reducing convergence rounds by more than half on DeepSeek-MoE-16B, and maintains superior performance under extreme sparsity by preserving critical synergistic expert interactions.
This paper addresses a core deployment bottleneck for Mixture-of-Experts (MoE) models: although sparse activation lets MoE architectures scale model capacity efficiently during training, the total memory footprint of all experts can remain prohibitive for serving. The authors argue that simply extracting a small subset of experts is not well solved by conventional Top-k selection heuristics, which score experts independently and then retain the highest-ranked ones. That pointwise view can miss two important phenomena: redundancy, where many high-scoring experts provide overlapping capabilities, and synergy, where certain experts are more valuable in combination than in isolation.
The key contribution is to reformulate expert selection as a dependency-aware set-selection problem rather than an independent ranking problem. The proposed method models the utility of an expert subset with a submodular objective that captures diminishing returns from redundant experts while allowing complementary experts to contribute additional value when selected together. By casting the problem in submodular optimization, the work enables a principled selection procedure—typically amenable to greedy or approximate algorithms—under an expert-budget constraint. The central insight is that the best compact MoE is not necessarily the one containing the individually strongest experts, but the one whose experts jointly provide diverse, non-redundant, and mutually reinforcing coverage.
This matters because it offers a more structured path toward efficient MoE deployment, especially when memory and latency constraints force practitioners to keep only a fraction of the original expert pool. Instead of relying on ad hoc importance scores, the approach provides a framework for balancing model capacity, expert diversity, and expert interactions. If effective, such dependency-aware selection can improve the quality–efficiency tradeoff of compressed MoE models, making sparse expert architectures more practical for real-world deployment while preserving the benefits of specialized expert computation.