AI EncyclopediaArchitectures & models
Mixture of Experts (MoE)
A model architecture that splits computation across many specialised sub-networks ('experts') and, for each input, routes it to only a few of them - giving a model the knowledge capacity of a huge network at the compute cost of a much smaller one.
In plain English
Instead of one giant network doing everything for every request, an MoE model is like a large team of specialists where only a couple are called in for each specific question. The whole team's combined expertise is available, but you only pay the compute cost of the two or three you actually used.
Technical explanation
A routing mechanism selects which expert sub-networks handle each input token, activating only a small subset while the rest stay idle - trading more memory (all experts must be stored) for less compute per token. Google's Switch Transformer research demonstrated up to 7x faster pre-training than an equivalent dense model at the same compute budget, establishing the technique's efficiency case.
Why it matters
MoE is why current frontier models can have enormous total parameter counts without a proportionally enormous inference cost. By 2026, nearly all frontier language models use some form of MoE routing rather than a single dense network.
Real-world example
Mistral AI's Mixtral 8x7B has 8 experts per layer with 2 active per token: 47 billion total parameters, but only about 13 billion active per token, letting it match or outperform the dense Llama 2 70B on most benchmarks at a fraction of the per-token compute.
Common misunderstanding
That MoE models are 'smaller' than their total parameter count suggests. They still require enough memory to hold every expert, even though only a few are active on any given token - the saving is in compute, not in the hardware footprint needed to serve the model.
Something wrong here?
Every entry is hand-researched and hand-written by Nakoda. If a fact is stale, a source has changed or a definition needs sharpening, tell us and we will check it.

