Md. Asif Uddin

    LectionesPart IX

    Sparsity and scale

    Why a model can have 671 billion parameters and use 37 billion of them.

    Every parameter count you read in a press release is now ambiguous, and this part is how it became ambiguous.

    The first two papers establish conditional computation and then simplify it until it trains reliably. The last three are the open models — the point at which a technique stops being a research direction and becomes something you download.

    Read the load-balancing losses closely. Routing collapses without them, everyone reinvents them, and each of these papers handles the problem slightly differently. DeepSeek-V3 removes the auxiliary loss altogether, which is the sort of claim worth checking carefully.

    The reading

    1. Sparse MoE

      Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

      Shazeer et al. · ICLR · 2017

      Claim
      A gating network that routes each example to a few of thousands of experts gives a thousand times more parameters at almost constant computation.
      Why
      Where conditional computation became practical. The load-balancing loss in Section 4 is the part everyone rediscovers the hard way.
      Read
      Sections 2 to 4.
    2. Switch

      Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity

      Fedus, Zoph & Shazeer · JMLR · 2021

      Claim
      Routing each token to exactly one expert is simpler, more stable and faster than the two that had been assumed necessary.
      Why
      The simplification that made mixture of experts trainable at scale, and the most honest published treatment of training instability on this list.
      Read
      Sections 2 and 5.
    3. Mixtral

      Mixtral of Experts

      Jiang et al. · Mistral AI · 2024

      Claim
      Eight experts with two active per token: 47 billion parameters in total, 13 billion used per token, matching or beating a dense 70-billion model.
      Why
      The first open mixture of experts anyone could run, and where the distinction between total and active parameters became something you have to say out loud.
    4. DeepSeekMoE

      DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

      Dai et al. · ACL · 2024

      Claim
      Splitting experts finer and reserving a few as always-on shared experts produces the specialisation that coarse routing never achieved.
      Why
      An argument about what an expert is for, supported by ablation rather than assertion. Read it as the counter-position to Switch.
      Read
      Section 3.
    5. DeepSeek-V3

      DeepSeek-V3 Technical Report

      DeepSeek-AI · Technical report · 2024

      Claim
      671 billion parameters with 37 billion active, trained in under 2.8 million H800-hours, with auxiliary-loss-free load balancing and multi-token prediction.
      Why
      The most detailed public account of training a frontier model that exists. Read it as an engineering document; the cost table is the argument and the infrastructure section is where the work is.
      Read
      Sections 3 and 4.