paper · Mrinank Sharma · tradition: machine-learning

Measures sycophancy: RLHF-tuned models systematically shift answers toward what the user appears to want rather than toward accuracy. Evidence that the reward favors the agreeable answer over the faithful one.

The book’s stance. cited in Ch 11 as evidence for objective capture (helpful/agreeable over faithful/corrective).

Availability. Cited by reference; no local copy held in the repo.

Where this is cited in the book