EVERYTHING AIAI engineering, made visual
Visual edition planned
Phase 06 · Lesson 15LearnPython1.3 h17 lessons in phase

Streaming Speech-to-Speech — Moshi, Hibiki, and Full-Duplex Dialogue

2024-2026 redefined voice AI. Moshi ships a single model that listens and speaks simultaneously at 200 ms latency. Hibiki does speech-to-speech translation chunk-by-chunk. Both abandon the ASR → LLM → TTS pipeline for a unified full-duplex architecture over Mimi codec tokens. This is the new reference design.

Visual edition planned

This lesson isn’t interactive yet.

It is part of the curriculum and will get the same treatment as Phase 1 — a visual cover, hands-on labs, derivations with numeric checks and a quiz. Until then, the original lesson is the best place to read it:

Part of Phase 06Speech & Audio. Use the previous / next cards below to keep browsing the phase.