Section 2
Training and Alignment
Section 1 built the mechanism that lets a Transformer process a sequence. This section follows several training stages used in modern model development: pretraining on next-token prediction, supervised or parameter-efficient adaptation, preference modeling, policy optimization, and evaluation. Classical RLHF is one path through those stages; direct preference methods and other post-training recipes use different components.
Chapters
4
Lesson Reading
1 hr 40 min
Labs
4
Interview Sets
4