Skip to content
mlmentorship

Book VII · Chapter 1

RL foundations and value methods

Start with MDPs, then learn value-based control, exploration, and reward design.

Priority
Specialist
Difficulty
Intermediate
Useful for
RS, RE, Post-training, Robotics
Interview rounds
ML breadth, Research

Read first: Probability and statistics

Chapter contents

6 entries · read in order
  1. 01
    Markov decision processes and Bellman equations
    Concept
  2. 02
    Value-based vs. policy-based RL
    Concept
  3. 03
    Q-learning
    Concept
  4. 04
    Exploration vs exploitation: epsilon-greedy, UCB, Thompson sampling
    Concept
  5. 05
    Contextual bandits
    Concept
  6. 06
    Reward shaping
    Concept