Skip to content
mlmentorship

Book V · Chapter 1

Hardware and performance

Count work and memory, understand accelerator limits, then read a trace.

Priority
Role-specific
Difficulty
Advanced
Useful for
RE, Systems MLE, Performance
Interview rounds
Systems, Performance

Read first: Neural network foundations

Chapter contents

5 entries · read in order
  1. 01
    Transformer compute and memory accounting
    Concept
  2. 02
    GPU memory hierarchy: HBM, SRAM, and roofline reasoning
    Concept
  3. 03
    Accelerator network topology for distributed ML
    Concept
  4. 04
    Profiling distributed ML workloads
    Concept
  5. 05
    Optimize an accelerator workload from a trace
    Question