Skip to content
mlmentorship

Book VIII

Vision, language, and speech

Visual models, multimodal systems, sequence modeling, natural language, and speech.

4 chapters · 21 entries

Chapters

4
VIII.1

Vision and multimodal foundations

Build from convolution to residual and transformer models, then study transfer and robustness.

Scope
Specialist
Difficulty
Intermediate
Useful for
Vision, Multimodal, AS, RS
VIII.2

Vision tasks

Move from object detection and suppression to dense semantic prediction.

Scope
Specialist
Difficulty
Intermediate
Useful for
Vision, Multimodal
VIII.3

Language representations and sequence models

Follow text representation from embeddings and recurrence to bidirectional encoders and structured prediction.

Scope
Specialist
Difficulty
Intermediate
Useful for
NLP, Speech, AS, RS
VIII.4

Speech and real-time multimodal systems

Build from ASR objectives and streaming speech to a complete live multimodal assistant.

Scope
Specialist
Difficulty
Advanced
Useful for
Speech, Multimodal, RS, RE, MLE