EVERYTHING AIAI engineering, made visual
Visual edition planned
Phase 09 · Lesson 09BuildPython45 min12 lessons in phase

Reward Modeling & RLHF

Humans cannot write a reward function for "good assistant response," but they can compare two responses and pick the better one. Fit a reward model to those comparisons, then RL the language model against it. Christiano 2017. InstructGPT 2022. The recipe that turned GPT-3 into ChatGPT. In 2026 it is mostly being replaced by DPO — but the mental model stays.

Visual edition planned

This lesson isn’t interactive yet.

It is part of the curriculum and will get the same treatment as Phase 1 — a visual cover, hands-on labs, derivations with numeric checks and a quiz. Until then, the original lesson is the best place to read it:

Part of Phase 09Reinforcement Learning. Use the previous / next cards below to keep browsing the phase.