Variable capacity
Evaluation and data collection surge around model launches, red-team exercises, and benchmark refreshes. A fixed internal team is sized for the trough or the peak, rarely both.
Data annotation outsourcing
Outsourcing simple annotation is easy. Outsourcing work that requires professional judgment is not. Medical AI may require clinicians. Legal AI may require attorneys. Financial AI may require finance professionals. Scientific models may require researchers and PhDs.
We are not a freelancer marketplace or traditional staffing company.
Data annotation outsourcing is buying managed external capacity for labeling, review, or evaluation so the buyer does not recruit and manage each professional. For expert AI work, the object is a qualified cohort that can apply a professional standard, not a public freelancer listing.
Data annotation outsourcing is the transfer of labeling, review, or evaluation work to an external team that the buyer does not have to recruit and manage person by person. For commodity tasks, that team can be a general workforce. For professional domains, it needs to be a qualified cohort: people who can apply a clinical, legal, financial, scientific, or engineering standard to data that looks plausible to a non-specialist.
Stratum provides managed professional expert capacity for projects where generic crowd annotation is insufficient. The buyer keeps the methodology. We provide the people and the operating layer around them.
Outsourcing is usually a response to capacity, scarcity, or time—not a claim that internal teams are incapable. The programs that benefit most have a clear expert profile and a reason the work should not sit on a small in-house queue forever.
Evaluation and data collection surge around model launches, red-team exercises, and benchmark refreshes. A fixed internal team is sized for the trough or the peak, rarely both.
The next domain is not “more of the same labels.” It is oncology after primary care, tax after commercial contracts, or repository-level coding after snippet review.
A six-week gold-set build does not justify a hiring plan. It does justify a defined cohort with a start date and an exit.
Held-out tasks need people who can write items that are difficult for the model and still scoreable by a professional standard.
Side-by-side model comparisons and regression suites consume reviewer hours faster than most research teams can spare.
A legal or clinical product that enters a new jurisdiction often needs reviewers who actually practice there, not a translated rubric.
Medicine, law, finance, science, engineering, and software produce errors that can look complete to a general reviewer. A discharge summary can omit a contraindication and still read fluently. A contract clause can be locally invalid and still look standard. A valuation model can double-count cash and still format cleanly. A paper citation can fail to support the claim. Code can pass a docstring check and fail the actual algorithm.
If those mistakes would be obvious to a professional and invisible to a crowd worker, the project is no longer a labeling problem. It is an expert judgment problem.
| Criterion | In house | Managed outsourcing |
|---|---|---|
| Speed | Faster if the experts already sit on the team | Faster when the next specialty is not on payroll |
| Control | Highest if the same people own method and labels | High if the client keeps rubrics and acceptance |
| Capacity | Hard to surge without hiring | Designed for time-bounded or variable volume |
| Recruiting burden | Stays internal | Moved to the operating vendor |
| Criterion | How talent is sourced | Who manages workers | Client workload | Best use |
|---|---|---|---|---|
| Freelancer marketplace | Buyer browses profiles and hires one by one | Mostly the buyer | High: recruiting, contracting, replacement | One-off commodity tasks |
| Traditional staffing | Candidates are placed into jobs | Employer after hire | High if the need is project capacity, not a seat | Standing roles, not time-bounded evaluation |
| Managed expert capacity | Vendor sources against a written specification | Vendor operates the cohort | Lower: client keeps methodology and acceptance | Expert AI data and evaluation programs |
Domains below are the starting map, not a public inventory. Specialties are sourced according to the project requirements.
Physicians · Registered nurses · Pharmacists · Clinical specialists
Attorneys · Legal researchers · Practice area specialists
Investment bankers · Financial analysts · CPAs · Accounting professionals · Private equity professionals
PhDs · Researchers · Bioinformaticians · Chemists · Biologists
Mechanical engineers · Electrical engineers · Technical specialists
Senior developers · Software engineers · Technical reviewers
Annotation is often the search term. The actual statement of work is usually broader: the same experts may write gold answers, score model traces, or adjudicate disagreements.
The process is designed so the buyer does not have to become a recruiting desk. It is also designed so quality is not postponed until the end of a large batch.
Step 1
Capture the domain, credentials, geography, volume, duration, task type, data access model, and the quality bar the client will accept.
Step 2
Translate the brief into a sourcing profile. This is a specification, not a job post and not a public freelancer listing.
Step 3
Verify credentials where required and run domain tasks that resemble production work rather than generic attention checks.
Step 4
Align reviewers on the rubric, gold examples, and edge cases before the program takes volume.
Step 5
Place qualified experts into the client’s tools, confidentiality terms, and workflow conventions.
Step 6
Operate the agreed capacity: hours, coverage windows, and task routing against the live queue.
Step 7
Apply review layers, gold-set checks, and disagreement handling according to the quality plan.
Step 8
Add, replace, or reduce experts without forcing the buyer to restart individual recruiting.
Expert work still needs a quality system. Credentials get a reviewer into the program. They do not guarantee that two specialists will apply a new rubric the same way on day one.
Use dual review on items that are high impact, highly ambiguous, or used as gold.
Hide model identity or other reviewers’ labels when the comparison should not be influenced by source.
Measure inter-reviewer agreement as a process signal. Low agreement can mean a weak rubric, a mixed cohort, or a genuinely hard item.
Seed production with adjudicated examples and refresh them when the product or rubric changes.
Give reviewers a path to say “this is outside my specialty” instead of forcing a guess.
Resolve remaining disputes with a more senior or more specialized professional, and feed the decision back into the rubric.
Security claims should match the operating model. Stratum can work inside client-controlled tools and environments where that is the safer design. Access is intended to be project-specific, least-privilege, and removed when the engagement ends.
This page does not claim SOC 2, HIPAA, ISO, or other certifications. Project-specific security requirements should be discussed before work starts. See the security page for the operating principles.
There is no honest single price for expert annotation. A licensed specialist reviewing high-ambiguity items with adjudication is a different product from a general labeling queue. Cost is a function of the specification.
Before sourcing experts, define the specification the project will be staffed and measured against.
The useful question is not “Who can provide annotators?” It is “Who can produce a stable expert cohort against this profile, with a quality system the research team can inspect?”
Share the profession, specialty, experience, location, headcount, hours, duration, and project description. We assemble the expert capacity.