LectionesPart IX
Sparsity and scale
Why a model can have 671 billion parameters and use 37 billion of them.
Every parameter count you read in a press release is now ambiguous, and this part is how it became ambiguous.
The first two papers establish conditional computation and then simplify it until it trains reliably. The last three are the open models — the point at which a technique stops being a research direction and becomes something you download.
Read the load-balancing losses closely. Routing collapses without them, everyone reinvents them, and each of these papers handles the problem slightly differently. DeepSeek-V3 removes the auxiliary loss altogether, which is the sort of claim worth checking carefully.
The reading
- Sparse MoE
Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Shazeer et al. · ICLR · 2017
- Claim
- A gating network that routes each example to a few of thousands of experts gives a thousand times more parameters at almost constant computation.
- Why
- Where conditional computation became practical. The load-balancing loss in Section 4 is the part everyone rediscovers the hard way.
- Read
- Sections 2 to 4.
- Switch
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
Fedus, Zoph & Shazeer · JMLR · 2021
- Claim
- Routing each token to exactly one expert is simpler, more stable and faster than the two that had been assumed necessary.
- Why
- The simplification that made mixture of experts trainable at scale, and the most honest published treatment of training instability on this list.
- Read
- Sections 2 and 5.
- Mixtral
Mixtral of Experts
Jiang et al. · Mistral AI · 2024
- Claim
- Eight experts with two active per token: 47 billion parameters in total, 13 billion used per token, matching or beating a dense 70-billion model.
- Why
- The first open mixture of experts anyone could run, and where the distinction between total and active parameters became something you have to say out loud.
- DeepSeekMoE
DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
Dai et al. · ACL · 2024
- Claim
- Splitting experts finer and reserving a few as always-on shared experts produces the specialisation that coarse routing never achieved.
- Why
- An argument about what an expert is for, supported by ablation rather than assertion. Read it as the counter-position to Switch.
- Read
- Section 3.
- DeepSeek-V3
DeepSeek-V3 Technical Report
DeepSeek-AI · Technical report · 2024
- Claim
- 671 billion parameters with 37 billion active, trained in under 2.8 million H800-hours, with auxiliary-loss-free load balancing and multi-token prediction.
- Why
- The most detailed public account of training a frontier model that exists. Read it as an engineering document; the cost table is the argument and the infrastructure section is where the work is.
- Read
- Sections 3 and 4.