Stratum

Insight

How to Build Gold-Standard AI Evaluation Data

A process for building gold-standard evaluation data: task specification, independent labeling, adjudication, maintenance, and recalibration.

Published 2026-08-25 · Written by the Stratum editorial team. No individual author or outside clinical or legal reviewer is named for this article.

Gold-standard AI evaluation data is a maintained set of items whose accepted labels have survived independent expert review and, where needed, adjudication. It is used to calibrate reviewers, measure models, and detect drift. It is not a pile of leftover production labels with the word “gold” added in a spreadsheet.

Task specification

Write what the item is, what the rater may look at, and what “correct” means in operational language. If two competent experts would need a meeting to interpret the instruction, the gold set will encode that meeting’s absences.

Expert selection

Use people who match the standard you want to measure. A gold set for specialty oncology scored by generalist raters is a generalist set. The Expert Necessity Test is the filter.

Independent labeling

At least two qualified reviewers should label without seeing each other’s answers when the item will be used as gold. Independence is what makes later agreement informative. Sequential “review the first person’s work” is a different, still useful, process.

Disagreement and adjudication

Expected disagreement is not a reason to drop the item. It is a reason to adjudicate and to write down why. Items that remain irresolvable may be poor gold candidates; they can still be useful as discussion cases.

Gold set maintenance

Assign an owner. Record the rubric version. Retire items when the product policy changes. Do not quietly edit a gold answer while leaving the old model scores in a dashboard as if they were comparable.

Drift and recalibration

Gold sets drift when the model family changes what “typical error” looks like, when reviewers turn over, and when the product adds tools. Recalibrate on a schedule and after material changes. A stale gold set creates false confidence: the numbers move, the standard does not.

Expert Evaluation Quality Loop

Quality is treated as a cycle, not a single inspection step at the end of a project.

  1. 01

    Qualification

  2. 02

    Calibration

  3. 03

    Production

  4. 04

    Review

  5. 05

    Disagreement detection

  6. 06

    Adjudication

  7. 07

    Feedback

  8. 08

    Recalibration

Related service pages: expert training data and quality control.

Need the capacity behind this process?

If the article describes work you are scoping now, send the expert specification rather than a general RFP.

  • Need 25 licensed nurses for a clinical AI evaluation
  • Need 15 attorneys in a specific practice area
  • Need 20 PhD scientists for benchmark creation
  • Need 30 senior software engineers for code evaluation