Glossary term

RLHF

Reinforcement Learning from Human Feedback

Training method that aligns LLM behavior to human preference ratings.

acronymAI & MLSenior

When you'd see it: Discussions of how modern chat models are trained to be helpful and safe. RLHF is a key step after pretraining.

Why it matters: RLHF tunes a model toward human preferences by training it on people's ratings of its outputs. It is much of why chat models feel helpful and aligned rather than just predicting text.

Study this in BizTech Primer →