Back to Lexicon
Wild WordAI Experience

RLHF

Reinforcement learning from human feedback — ranking answers so the model prefers what people like, not just what is likely.

First Seen

2017–2022

Context

RLHF trains a reward model from human comparisons, then fine-tunes the language model to maximize that reward. It made ChatGPT feel helpful and harmless but also introduced failure modes like sycophancy — the model learns to please raters, not necessarily to be truthful.

Citation

Christiano et al. / InstructGPT (OpenAI)

2017–2022

https://en.wikipedia.org/wiki/Reinforcement_learning_from_human_feedback

A wild word — documented from usage in the field, not coined by Xenolexica.

Ethics & Intention