New: conversational Urdu and Punjabi speech corpora now licensable  See the catalog →

Practical guides

How to actually measure data annotation quality

Gold seeding, double-blind passes and senior adjudication — the three layers that produce trustworthy labels, and what a delivery report must contain.

Key takeaways

  • One quality number with no methodology behind it tells you nothing. Ask how it was computed.
  • Three independent layers work: gold seeding, double-blind passes, and senior adjudication.
  • Gold questions must be invisible and continuous, not a one-off entry test.
  • The most useful thing in a delivery report is the list of items the vendor rejected before you saw them.

Why a single accuracy number is meaningless

“99% accuracy” is the most common claim in this industry and among the least informative. Accuracy against what reference? Measured on what sample? Adjudicated by whom? Computed before or after the vendor's own filtering?

Without those answers the number is marketing. A vendor who explains the methodology unprompted is telling you something real; one who quotes the figure and moves on is not.

Layer one: gold seeding

Pre-adjudicated items are injected invisibly into live annotation queues at a set rate. Each annotator carries a rolling accuracy score computed against them.

Two properties matter. Invisible: if annotators can identify gold items they will treat them differently and the measurement stops being representative. Continuous: a one-off qualification test measures entry competence, not performance in week six under production pressure. When a score drops below threshold the queue should pause automatically rather than degrading quietly.

Layer two: double-blind passes

A defined proportion of items — 10% as a baseline, up to 100% for high-stakes schemas — get labelled independently by two annotators who cannot see each other's work.

Agreement is then computed per batch and, critically, per label class. The aggregate number tells you whether things are broadly fine. The per-class breakdown tells you which categories are genuinely ambiguous, which is the information you can actually act on.

Layer three: senior adjudication

Disagreements above threshold route to a senior linguist or domain expert whose decision is final and, importantly, feeds back into the written rubric.

The feedback loop is what makes this a system rather than a cleanup step. Rubric changes should be versioned, so when a labelling convention shifts in week six you can trace why and re-examine earlier batches if needed.

What a delivery report must contain

Per-annotator accuracy distribution — the spread, not just the mean. Inter-annotator agreement by label class. The adjudication log. Throughput against plan. The rubric version in force. And the list of items rejected internally before delivery.

That final item is the strongest signal available. A vendor willing to show you their rejects is showing you their standard. One reporting only successes is asking for trust with no evidence attached.

Fraud is a permanent condition, not an incident

At scale, some fraction of submissions will be fraudulent: one person operating several accounts, recycled work, model output submitted as human labour. This is a constant industry pressure and a vendor who seems surprised by it has not operated at volume.

The defences that work are structural: identity verification at onboarding, device and location fingerprinting, duplicate fingerprinting across the whole corpus, attention checks embedded in tasks, and payment tied to QA-passed units rather than submissions.

Related pages on this site

Answers

Frequently asked questions

Injecting pre-adjudicated items invisibly into live annotation queues to measure each annotator's accuracy continuously. It differs from a qualification test, which measures competence at entry rather than performance under sustained production conditions.

10% double-blind is a reasonable baseline for stable schemas. High-stakes work — safety classification, medical, legal — often warrants 100%. The right figure depends on the cost of an undetected error, which is a decision only you can make.

Ask for the delivery report and read past the headline. Look for per-annotator accuracy distribution, agreement by label class, an adjudication log, a rubric version, and the items they rejected internally. A vendor who cannot produce those is reporting a number rather than running a system.

Send us 500 items and your schema.

We will label them free, return the accuracy and agreement numbers, and show you the disagreements.

Average first response: under 6 business hours. NDAs signed same day.

Free samples Get a quote