EVERYTHING AIAI engineering, made visual
Visual edition planned
Phase 10 · Lesson 08Build1.3 h24 lessons in phase

DPO: Direct Preference Optimization

RLHF works. It also requires training three models (SFT, reward model, policy), managing PPO's instability, and tuning a KL penalty. DPO asks: what if you could skip all of that? DPO directly optimizes the language model on preference pairs. No reward model. No PPO. One training loop. Same results.

Visual edition planned

This lesson isn’t interactive yet.

It is part of the curriculum and will get the same treatment as Phase 1 — a visual cover, hands-on labs, derivations with numeric checks and a quiz. Until then, the original lesson is the best place to read it:

Part of Phase 10LLMs from Scratch. Use the previous / next cards below to keep browsing the phase.