EVERYTHING AIAI engineering, made visual
Visual edition planned
Phase 06 · Lesson 04BuildPython45 min17 lessons in phase

Speech Recognition (ASR) — CTC, RNN-T, Attention

Speech recognition is audio classification at every timestep, glued together by a sequence model that knows English and silence. CTC, RNN-T, and attention are the three ways to do it. Pick one and understand why.

Visual edition planned

This lesson isn’t interactive yet.

It is part of the curriculum and will get the same treatment as Phase 1 — a visual cover, hands-on labs, derivations with numeric checks and a quiz. Until then, the original lesson is the best place to read it:

Part of Phase 06Speech & Audio. Use the previous / next cards below to keep browsing the phase.