Definition
What is RLHF
RLHF is reinforcement learning from human feedback. Humans provide preferences, critiques, or other signals that are used to train a reward model or to update a policy. This page describes the human-data role, not a claim about a specific training stack.
What is rlhf?
RLHF is reinforcement learning from human feedback. Humans provide preferences, critiques, or other signals that are used to train a reward model or to update a policy. This page describes the human-data role, not a claim about a specific training stack.
Why it matters
The quality of RLHF is bounded by who provides the feedback and how disagreement is handled.
Example
Experts rank safety-critical answers; those labels train or evaluate a reward model. See also OpenAI’s public writing on learning from human feedback.
When it is used
The program needs a human preference signal for post-training, not only SFT pairs.
Common mistakes
- Treating RLHF as a synonym for any annotation
- Unqualified raters on professional domains
- No held-out expert evaluation after training
Sources
Citations point to primary technical or institutional documents. They support definitions, not a claim that those organizations are customers.
- OpenAI: Learning from human preferences — Primary description of using human preference comparisons in model training.
- Bai et al., Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback — Peer-available Anthropic research on RLHF as a training signal.
What kind of experts do you need?
Share the profession, specialty, experience, location, headcount, hours, duration, and project description. We assemble the expert capacity.
- Need 25 licensed nurses for a clinical AI evaluation
- Need 15 attorneys in a specific practice area
- Need 20 PhD scientists for benchmark creation
- Need 30 senior software engineers for code evaluation