Definitions for expert AI data work
Short definitions for the concepts that sit under expert annotation, evaluation, and post-training. Each page is written to stand alone if a search or retrieval system extracts it.
Expert data annotation
Expert data annotation is the labeling, scoring, or authoring of AI data by people who can apply a professional standard. It is used when an incorrect answer could look plausible to a general reviewer but obvious to a practitioner.
Gold dataset
A gold dataset is a maintained set of items whose accepted labels have survived independent expert review and, where needed, adjudication. It is used to calibrate reviewers, measure models, and detect drift.
Expert adjudication
Expert adjudication is senior review that resolves disagreement between qualified reviewers and records a final accepted label or score, including a reason that can change the guide.
Inter-annotator agreement
Inter-annotator agreement is a measure of how often two or more reviewers assign the same label to the same item. Simple percent agreement is a starting diagnostic; chance-corrected coefficients exist for more formal analysis.
AI evaluation rubric
An AI evaluation rubric is the scoring contract for a program: the dimensions, examples, severity scale, and decision rules that make reviewer labels comparable over time.
Preference data
Preference data is pairwise or ranked human judgment about which model output should be preferred. For expert programs, the rater should be able to justify the preference in domain terms.
RLHF
RLHF is reinforcement learning from human feedback. Humans provide preferences, critiques, or other signals that are used to train a reward model or to update a policy. This page describes the human-data role, not a claim about a specific training stack.
SFT data
SFT data is supervised fine-tuning data: prompt and target-response pairs used to teach a model a behavior, format, or domain style. Expert SFT is authored or heavily edited by practitioners.
AI benchmark
An AI benchmark is a held-out collection of tasks and scoring rules used to measure model capability over time. A useful expert benchmark is representative of real failures and is not used as training data for the same measurement.
Domain expert evaluation
Domain expert evaluation is professional review of model outputs by people who already practice the field the system claims to serve. It asks whether the output is professionally acceptable, not only whether it sounds better than an alternative.
What kind of experts do you need?
Share the profession, specialty, experience, location, headcount, hours, duration, and project description. We assemble the expert capacity.
- Need 25 licensed nurses for a clinical AI evaluation
- Need 15 attorneys in a specific practice area
- Need 20 PhD scientists for benchmark creation
- Need 30 senior software engineers for code evaluation