Models · reviewed · reviewed Oct 5, 2026 · 4 min
How does a Mixture of Experts model work?
A sparse Mixture of Experts layer routes a token representation to selected learned blocks and combines their outputs. Other blocks stay available but do not run for that token. Total parameters describe stored capacity; active parameters include the selected experts and shared layers.
One token, several possible paths
A transformer carries a numerical representation of each token through repeated layers. Attention mixes information between permitted positions; a feed-forward block transforms each position’s representation. In a common sparse Mixture of Experts, or MoE, design, selected feed-forward blocks are replaced by a bank of alternatives.
An expert is a learned block of parameters. A router scores the alternatives for the current token representation. It selects a small subset, sends the representation through those blocks, and combines their outputs. The resulting representation continues through the same network.
This is conditional computation: the input helps determine which parameters participate. It does not create a conversation between independent assistants. An expert has no separate chat history, tool permissions, or harness loop.
Watch selection change the calculation
Suppose a tiny layer has three experts. Its router sees a two-number input, chooses the highest-scoring expert or pair, and assigns mixing weights within that selected set. With one selected expert, that expert gets the entire mixing weight. With two, both outputs can contribute.
The experiment replaces real learned networks with three simple authored functions. Change input A to B or C and watch the selected blocks change. Then switch to top-2 routing: predict whether the available count changes, and whether an inactive block can contribute to the output.
Explore the mechanism
Route a token through selected experts
[1, 0.2]↓ router scores → select top 1 → normalize selected gates- E1Selected
2x + y- Router probability
- 73.8%
- Mixing weight
- 100.0%
- Expert output
- 2.200
- E2Inactive
x − 2y- Router probability
- 22.2%
- Mixing weight
- 0.0%
- Expert output
- Not executed
- E3Inactive
−x + y- Router probability
- 4.0%
- Mixing weight
- 0.0%
- Expert output
- Not executed
- Available expert blocks
- 3
- Active expert blocks
- 1
Changing the route changes this token’s computation. It does not remove the other expert blocks from the network.
Inputs, router coefficients, and expert functions are authored teaching data. Probabilities and outputs are computed locally. Three scalar functions stand in for learned feed-forward blocks; no LLM runs, and block counts are not memory or latency measurements.
The three router probabilities describe the initial comparison. Mixing weights describe the subsequent combination: inactive experts have zero weight, and selected weights sum to one. These are two different stages, so a probability of 20% does not imply that an unselected expert contributes 20% of the output.
Input A under top-1 routing runs E1 and produces 2.2. Top-2 routing includes E2 as well; the combined value changes because the selected outputs differ. These numbers are locally computed consequences of the disclosed functions, not evidence of an LLM’s accuracy.
Stored capacity and active computation
All three experts still belong to the toy layer when only one runs. Real sparse models similarly distinguish total parameters from active parameters. Both descriptions also need to account for shared parts of the network, such as attention and embeddings.
The original Mixtral architecture routes each token through two of eight feed-forward experts per MoE layer. Switch Transformer explores top-1 routing. These are concrete designs, not rules that every MoE must use the same expert count or gate.
Increasing the expert bank while keeping the selected count fixed can add parameter capacity without multiplying the expert computation for each token by the same factor. It still adds weights that must be stored or made accessible. “Only some parameters are active” does not mean “only those parameters need memory.”
The demo’s block counter deliberately excludes shared computation. It cannot be converted directly into model size, token latency, throughput, or GPU requirements.
The router has to serve a whole batch
Imagine ten token representations all choosing E1. E2 and E3 may sit idle while E1 receives most of the work. A routing decision that is sensible for one token can produce an awkward workload across a batch.
Sparse MoE research therefore addresses load balancing, communication between devices, and efficient execution of uneven expert batches. Switch Transformer uses an expert capacity policy; excess routed tokens can skip that expert transformation through the residual path. Other implementations organize capacity differently.
More experts are not a universal speedup. The actual result depends on routing, shared work, batching, hardware placement, memory access, and execution kernels. Count active parameters to understand the architecture; measure the deployed workload to understand its performance.
“Expert” does not guarantee a human specialty
E1 is named E1 because assigning it a label such as “coding expert” would invent a meaning the experiment does not establish. Learned routing may reflect patterns that do not align neatly with human subject categories. Mixtral’s routing analysis does not find a simple one-expert-per-domain organization.
Keep three different ideas separate. MoE chooses blocks inside a network. An ensemble combines predictions from multiple models. Harness orchestration coordinates separate agent runs and their effects. Distillation can train a student from a teacher’s signals; quantization changes numerical representations. None of those mechanisms follows automatically from the word “experts.”
Inspect the toy router and combination rule
The input representations are A [1, 0.2], B [0.1, 1], and C [−0.8, −0.6]. For input [x, y], router scores are 1.8x − 0.4y, 0.2x + 1.6y, and −x − y.
Softmax converts scores to positive probabilities summing to one. The top one or two scores are selected. Selected probabilities are divided by their selected sum; other gates become zero. This is equivalent here to softmax over the selected scores.
Expert outputs are E1 2x + y, E2 x − 2y, and E3 −x + y. The combined output is the sum of each selected output multiplied by its gate. Inactive functions are not executed. There is no training, capacity overflow, attention, or device communication in this widget.
Sources
Sources and further reading
- 01Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts LayerShazeer et al. · research · published Jan 23, 2017 · source checked Oct 5, 2026
Primary mechanism for sparse learned gating, selected feed-forward experts, load balancing, and communication costs.
- 02Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient SparsityFedus, Zoph, and Shazeer · research · published Jan 11, 2021 · source checked Oct 5, 2026
Top-one expert routing and capacity limits distinguish per-token sparse computation from stored model capacity.
- 03Mixtral of ExpertsJiang et al. · research · published Jan 8, 2024 · source checked Oct 5, 2026
Concrete transformer MoE architecture with top-two normalized gates; explains active versus total parameters and analyzes routing patterns.
