Book IV · Chapter 4
Post-training, alignment, and safety
Preference learning, verifiable rewards, oversight, threats, and red-team design.
Read first: Reinforcement learning foundations
Chapter contents
9 entries · read in order- 01 RLHF, DPO, and the alignment training stack✓ Concept
- 02 Preference data and reward models✓ Concept
- 03 RL with verifiable rewards and GRPO✓ Concept
- 04 Scalable oversight and AI control✓ Concept
- 05 Model organisms of misalignment✓ Concept
- 06 LLM security threat models✓ Concept
- 07 Design post-training data, an RL environment, and its grader✓ Question
- 08 Design an LLM red-team and security evaluation program✓ Question
- 09 Design a safety control plane for high-impact agents✓ Question