EVERYTHING AIAI engineering, made visual
Visual edition planned
Phase 09 · Lesson 04BuildPython1.3 h12 lessons in phase

Temporal Difference — Q-Learning & SARSA

Monte Carlo waits until the episode ends. TD updates after every step by bootstrapping the next value estimate. Q-learning is off-policy and optimistic; SARSA is on-policy and cautious. Both are one line of code. Both underpin every deep-RL method in this phase.

Visual edition planned

This lesson isn’t interactive yet.

It is part of the curriculum and will get the same treatment as Phase 1 — a visual cover, hands-on labs, derivations with numeric checks and a quiz. Until then, the original lesson is the best place to read it:

Part of Phase 09Reinforcement Learning. Use the previous / next cards below to keep browsing the phase.