Book IV · Chapter 1
Transformer architecture and attention
Build from the transformer block to efficient and sparse attention variants.
Read first: Neural network foundations
Chapter contents
8 entries · read in order- 01 Transformer architecture: a senior-level mental model✓ Concept
- 02 Multi-head attention: why one head is not enough✓ Concept
- 03 Self-attention vs cross-attention✓ Concept
- 04 Grouped-query and multi-query attention (GQA, MQA)✓ Concept
- 05 FlashAttention✓ Concept
- 06 Sparse attention (BigBird, Longformer)✓ Concept
- 07 Linear attention (Linformer, Performer, kernel methods)✓ Concept
- 08 Mixture of Experts (MoE)✓ Concept