RLHF
Reinforcement Learning from Human Feedback
Training method that aligns LLM behavior to human preference ratings.
AI & ML
When you'd see it: Discussions of how modern chat models are trained to be helpful and safe. RLHF is a key step after pretraining.
Why it matters: RLHF tunes a model toward human preferences by training it on people's ratings of its outputs. It is much of why chat models feel helpful and aligned rather than just predicting text.
Study this in BizTech Primer →