Book VIII
Vision, language, and speech
Visual models, multimodal systems, sequence modeling, natural language, and speech.
4 chapters · 21 entriesChapters
4Vision and multimodal foundations
Build from convolution to residual and transformer models, then study transfer and robustness.
Vision tasks
Move from object detection and suppression to dense semantic prediction.
Language representations and sequence models
Follow text representation from embeddings and recurrence to bidirectional encoders and structured prediction.
Speech and real-time multimodal systems
Build from ASR objectives and streaming speech to a complete live multimodal assistant.