Stratum

Insight

How to Run an Expert AI Evaluation Pilot

A five-step pilot that tests the rubric and the cohort, with a written go/no-go, not a workshop.

Published 2026-08-26 · Written by the Stratum editorial team. No individual author or outside clinical or legal reviewer is named for this article.

An expert AI evaluation pilot tests the rubric and the cohort, not the sales process. Use real items, a written pass bar, and a planned decision: continue, revise the brief, or stop.

  1. Step 1

    Freeze a small item bank

    Twenty to one hundred items is a common planning range. Include easy, hard, and out-of-specialty cases.

  2. Step 2

    Qualify a thin cohort

    Enough people to produce disagreement, not a full production bench.

  3. Step 3

    Calibrate once

    Review gold examples together or asynchronously before scoring the bank.

  4. Step 4

    Score independently

    Hide other reviewers’ labels. Measure agreement and time per item.

  5. Step 5

    Adjudicate and decide

    Resolve disputes, update the rubric if needed, and write a go/no-go note.

Pilot checklist

  • Access path agreed before anyone sees data.
  • Pass bar written down.
  • Time-per-item captured for later pricing.
  • Disagreement treated as information.
  • No promise that pilot scores predict production SLAs.

Use Request a Pilot when the brief is ready, and the SLA article before you turn a pilot into a standing program.

Need the capacity behind this process?

If the article describes work you are scoping now, send the expert specification rather than a general RFP.

  • Need 25 licensed nurses for a clinical AI evaluation
  • Need 15 attorneys in a specific practice area
  • Need 20 PhD scientists for benchmark creation
  • Need 30 senior software engineers for code evaluation