EVERYTHING AIAI engineering, made visual
Visual edition planned
Phase 07 · Lesson 11BuildPython45 min16 lessons in phase

Mixture of Experts (MoE)

A dense 70B transformer activates every parameter for every token. A 671B MoE activates only 37B per token and beats it on every benchmark. Sparsity is the most important scaling idea of the decade.

Visual edition planned

This lesson isn’t interactive yet.

It is part of the curriculum and will get the same treatment as Phase 1 — a visual cover, hands-on labs, derivations with numeric checks and a quiz. Until then, the original lesson is the best place to read it:

Part of Phase 07Transformers Deep Dive. Use the previous / next cards below to keep browsing the phase.