Book II · Chapter 4
Efficiency, data, and scaling
Reduce memory and work, build training data, and reason about scale.
Read first: Numerics and training stability
Chapter contents
9 entries · read in order- 01 Activation checkpointing✓ Concept
- 02 Gradient accumulation✓ Concept
- 03 Sequence packing with block-diagonal masks✓ Concept
- 04 Foundation-model data curation✓ Concept
- 05 Synthetic data generation and verification✓ Concept
- 06 Design a foundation-model data platform✓ Question
- 07 Neural scaling laws and compute-optimal training✓ Concept
- 08 Knowledge distillation✓ Concept
- 09 Pruning: structured vs unstructured sparsity✓ Concept