EVERYTHING AIAI engineering, made visual
Visual edition planned
Phase 15 · Lesson 02Learn1 h22 lessons in phase

STaR, V-STaR, Quiet-STaR — Self-Taught Reasoning

The smallest possible self-improvement loop sits inside the rationale. A model generates a chain of thought, keeps the ones that land on correct answers, and fine-tunes on those. That is STaR. V-STaR adds a verifier so inference-time selection is better. Quiet-STaR pushes the rationale down to every token. All three work. None of them are magic — the loop preserves any shortcut that happened to reach the right answer.

Visual edition planned

This lesson isn’t interactive yet.

It is part of the curriculum and will get the same treatment as Phase 1 — a visual cover, hands-on labs, derivations with numeric checks and a quiz. Until then, the original lesson is the best place to read it:

Part of Phase 15Autonomous Systems. Use the previous / next cards below to keep browsing the phase.