Skip to content
mlmentorship

Book VIII · Chapter 4

Speech and real-time multimodal systems

Build from ASR objectives and streaming speech to a complete live multimodal assistant.

Priority
Specialist
Difficulty
Advanced
Useful for
Speech, Multimodal, RS, RE, MLE
Interview rounds
ML breadth, System design, Research

Read first: Language representations and sequence models

Chapter contents

7 entries · read in order
  1. 01
    Automatic speech recognition (ASR)
    Concept
  2. 02
    Connectionist Temporal Classification (CTC)
    Concept
  3. 03
    RNN-Transducer (RNN-T)
    Concept
  4. 04
    Streaming automatic speech recognition
    Concept
  5. 05
    Hybrid versus end-to-end speech recognition
    Concept
  6. 06
    Speaker recognition
    Concept
  7. 07
    Design a real-time multimodal assistant
    Question