Stratum

Definition

What is RLHF

RLHF is reinforcement learning from human feedback. Humans provide preferences, critiques, or other signals that are used to train a reward model or to update a policy. This page describes the human-data role, not a claim about a specific training stack.

What is rlhf?

RLHF is reinforcement learning from human feedback. Humans provide preferences, critiques, or other signals that are used to train a reward model or to update a policy. This page describes the human-data role, not a claim about a specific training stack.

Why it matters

The quality of RLHF is bounded by who provides the feedback and how disagreement is handled.

Example

Experts rank safety-critical answers; those labels train or evaluate a reward model. See also OpenAI’s public writing on learning from human feedback.

When it is used

The program needs a human preference signal for post-training, not only SFT pairs.

Common mistakes

  • Treating RLHF as a synonym for any annotation
  • Unqualified raters on professional domains
  • No held-out expert evaluation after training

Sources

Citations point to primary technical or institutional documents. They support definitions, not a claim that those organizations are customers.

What kind of experts do you need?

Share the profession, specialty, experience, location, headcount, hours, duration, and project description. We assemble the expert capacity.

  • Need 25 licensed nurses for a clinical AI evaluation
  • Need 15 attorneys in a specific practice area
  • Need 20 PhD scientists for benchmark creation
  • Need 30 senior software engineers for code evaluation