Wild WordAI Experience
RLHF
“Reinforcement learning from human feedback — ranking answers so the model prefers what people like, not just what is likely.”
First Seen
2017–2022
Context
RLHF trains a reward model from human comparisons, then fine-tunes the language model to maximize that reward. It made ChatGPT feel helpful and harmless but also introduced failure modes like sycophancy — the model learns to please raters, not necessarily to be truthful.
Citation
Christiano et al. / InstructGPT (OpenAI)
2017–2022
https://en.wikipedia.org/wiki/Reinforcement_learning_from_human_feedbackA wild word — documented from usage in the field, not coined by Xenolexica.
Ethics & Intention