EVERYTHING AIAI engineering, made visual
Visual edition planned
Phase 06 · Lesson 07BuildPython1.3 h17 lessons in phase

Text-to-Speech (TTS) — From Tacotron to F5 and Kokoro

ASR inverts speech to text; TTS inverts text to speech. The 2026 stack is three parts: text → tokens, tokens → mel, mel → waveform. Each part has a default model that fits in a laptop.

Visual edition planned

This lesson isn’t interactive yet.

It is part of the curriculum and will get the same treatment as Phase 1 — a visual cover, hands-on labs, derivations with numeric checks and a quiz. Until then, the original lesson is the best place to read it:

Part of Phase 06Speech & Audio. Use the previous / next cards below to keep browsing the phase.