paper · Long Ouyang · tradition: machine-learning

The InstructGPT paper: reinforcement learning from human feedback (RLHF) as the standard post-training method for aligning a model to human preference labels. In the book’s terms, the training objective as a selection criterion lifted into the architecture.

The book’s stance. cited in Ch 11 for RLHF as objective-level selection (objective capture).

Availability. Cited by reference; no local copy held in the repo.

Where this is cited in the book