Ilmu Komputer & AI editorial
You Only Need 2/3 of the Chosen Experts: An Empirical Study of Dynamic Expert Pruning in Fine-Grained MoE LLMs
The core problem
Fine-grained mixture-of-experts (MoE) architectures have become a mainstream design for open-weight large language models. Modern checkpoints now contain hundreds of experts and select an increasingly large number of them per token. This shift makes dynamic expert pruning an attractive route to cheaper inference, because computation at serving time scales with the number of experts actually executed for each token.
However, existing evidence on expert redundancy comes largely from coarser MoE architectures and from likelihood-scored multiple-choice benchmarks. It therefore remains unclear whether those findings transfer to the fine-grained regime. The authors identify three central open questions:
1. How redundant is per-token expert selection in fine-grained MoEs?
2. How effectively do existing dynamic pruning methods exploit that redundancy?
3. What governs a model's sensitivity to aggressive pruning?
To answer these questions, the paper conducts a systematic empirical study of **twelve fine-grained MoE checkpoints** spanning **nine architecture families**, evaluated on a core suite of **eleven benchmarks** covering knowledge QA, mathematics, code generation, and general reaso
Innovation
The empirical results are organized around the three research questions.
**Redundancy of expert selection.** Expert selection is far more redundant than the field's operating points assume. Uniformly retaining about **two thirds** of the selected experts preserves **98.8% of unpruned performance on average**. This requires only a one-integer change to the serving configuration and delivers **1.2-1.7x measured speedup across two serving backends**.
**Effectiveness of dynamic pruning at conservative budgets.** The simple uniform baseline leaves little room for dynamic allocation. At conservative budgets, even the best published rules differ from uniform truncation by **under 1%** at matched expert budgets.
**Effectiveness under aggressive pruning.** The value of dynamic rules emerges under aggressive pruning. There, the best rules recover **up to 3.0%** over uniform truncation. These gains are concentrated in the generative tasks that suffer the sharpest degradation when experts are removed.
**Model-level sensitivity.** Sensitivity to aggressive pruning depends on the model. Larger models and thinking models are more resilient, whereas multimodal models are more vulnerable.
| Fi
Why it matters
The study reframes the practical question for fine-grained MoE serving. Rather than asking which dynamic allocation rule is best, the more consequential question is how much expert computation a model can dispense with at all. The answer, on this evidence, is roughly one third of the selected experts per token, at a cost of only 1.2% average performance loss.
This has two implications. First, for deployment, the uniform two-thirds operating point is a strong default: it is trivial to implement, requires only a one-integer change, and captures most of the available savings. Second, for research, dynamic allocation earns its complexity only under aggressive pruning budgets, where generative tasks degrade most sharply and where the best rules recover up to 3.0% over uniform truncation.
The sensitivity results add a model-dependent dimension. Larger and thinking models tolerate aggressive pruning better, while multimodal models are more vulnerable. This suggests that pruning policy should be conditioned on model family rather than treated as a universal setting.
Together, these findings reveal how much expert computation fine-grained MoEs can dispense with, and establish when dynamic allocation earns its complexity, informing both practical deployment and future pruning methods.
Who should read this
Opening member contentโฆ