Skip to content
mlmentorship

Book VII

Reinforcement learning and robotics

Sequential decisions, value and policy methods, environments, rewards, and robotics policy learning.

3 chapters · 13 entries

Chapters

3
VII.1

RL foundations and value methods

Start with MDPs, then learn value-based control, exploration, and reward design.

Scope
Specialist
Difficulty
Intermediate
Useful for
RS, RE, Post-training, Robotics
VII.2

Policy optimization

Build policy gradients, actor-critic methods, advantage estimation, and PPO.

Scope
Specialist
Difficulty
Advanced
Useful for
RS, RE, Post-training, Robotics
VII.3

Environments, multiple agents, and robotics

Design environments and reason about interacting agents and learned policies in the physical world.

Scope
Specialist
Difficulty
Advanced
Useful for
RS, RE, Post-training, Robotics