Skip to content
NakodaAI

AI EncyclopediaTechniques & methods

Reinforcement Learning from Human Feedback (RLHF)

Also known as RLHF

A post-training method that adjusts a pre-trained model's behaviour using human preference judgments, rather than more raw text - the technique that turned raw language models into helpful assistants.

RLHFalignmentInstructGPTpost-training

In plain English

Pre-training teaches a model to predict text; it doesn't teach it to be helpful, honest or safe to talk to. RLHF is the extra step where humans rank the model's answers - this one is better than that one - and the model is nudged, over many rounds, toward the answers people actually prefer.

Technical explanation

RLHF is a three-stage process: (1) supervised fine-tuning on human-written demonstrations, priming the model to respond in the expected format; (2) reward-model training, where human raters rank or score model outputs and a separate model learns to predict those human preferences; (3) reinforcement-learning optimisation (commonly PPO), where the original model is updated using the reward model's scores on newly generated outputs. Research on this method found a 1.3-billion-parameter model with RLHF outperformed a 175-billion-parameter model without it on human preference evaluations.

Why it matters

RLHF is why a raw pre-trained LLM feels different from a released chat assistant - it is the alignment step that shapes helpfulness, tone and refusal behaviour, and it is central to current AI-safety practice at every major lab.

Real-world example

OpenAI's InstructGPT - the direct precursor to ChatGPT's approach - combined supervised fine-tuning on human demonstrations with PPO-based RLHF using a reward model trained from human rankings of outputs.

Common misunderstanding

That RLHF makes a model 'more accurate.' It optimises for what human raters prefer, which correlates with but is not identical to factual accuracy - a fluent, agreeable, well-formatted wrong answer can still score well with raters, which is one reason RLHF alone does not solve hallucination.

Something wrong here?

Every entry is hand-researched and hand-written by Nakoda. If a fact is stale, a source has changed or a definition needs sharpening, tell us and we will check it.