EVERYTHING AIAI engineering, made visual
Visual edition planned
Phase 09 · Lesson 06BuildPython1.3 h12 lessons in phase

Policy Gradient — REINFORCE from Scratch

Stop estimating value. Parameterize the policy directly, compute the gradient of expected return, step uphill. Williams (1992) wrote it in one theorem. It is why PPO, GRPO, and every LLM RL loop exist.

Visual edition planned

This lesson isn’t interactive yet.

It is part of the curriculum and will get the same treatment as Phase 1 — a visual cover, hands-on labs, derivations with numeric checks and a quiz. Until then, the original lesson is the best place to read it:

Part of Phase 09Reinforcement Learning. Use the previous / next cards below to keep browsing the phase.