Book VIII · Chapter 4
Speech and real-time multimodal systems
Build from ASR objectives and streaming speech to a complete live multimodal assistant.
Read first: Language representations and sequence models
Chapter contents
7 entries · read in order- 01 Automatic speech recognition (ASR)✓ Concept
- 02 Connectionist Temporal Classification (CTC)✓ Concept
- 03 RNN-Transducer (RNN-T)✓ Concept
- 04 Streaming automatic speech recognition✓ Concept
- 05 Hybrid versus end-to-end speech recognition✓ Concept
- 06 Speaker recognition✓ Concept
- 07 Design a real-time multimodal assistant✓ Question