New: conversational Urdu and Punjabi speech corpora now licensable  See the catalog →

Definitions

Inter-annotator agreement: what the number actually tells you

Cohen's kappa, Krippendorff's alpha and raw agreement — what each measures, what counts as acceptable, and why low agreement usually means bad guidelines.

Key takeaways

  • Raw percentage agreement is misleading because it ignores agreement that would occur by chance.
  • Cohen's kappa corrects for chance between two annotators; Krippendorff's alpha handles more annotators and missing data.
  • Above 0.8 is generally strong; 0.67–0.8 is usually workable; below 0.67 warrants investigation.
  • Low agreement almost always means ambiguous guidelines, not careless annotators.

Why raw agreement misleads

If two annotators label 1,000 items as spam or not-spam and 95% of items are not-spam, two people who both simply always answered not-spam would show 95% raw agreement while demonstrating no skill whatsoever.

Chance-corrected metrics exist to strip that out. They answer a better question: how much do these annotators agree beyond what random labelling at the observed base rates would produce.

Which metric to use

Cohen's kappa for exactly two annotators on categorical labels. Widely understood, easy to compute, and the default in most annotation tooling.

Fleiss's kappa generalises to more than two annotators where different items may be rated by different pairs.

Krippendorff's alpha is the most flexible: any number of annotators, missing data, and nominal, ordinal or interval scales. It is the right default for real annotation programmes where coverage is uneven.

For spans and bounding boxes, agreement is usually reported as F1 or IoU overlap rather than kappa, since the unit of agreement is a region rather than a category.

What counts as acceptable

Common convention: above 0.8 is strong, 0.67 to 0.8 is workable for most purposes, and below 0.67 means the results should not be trusted without investigation.

Treat those bands as guidance rather than law. A genuinely subjective task — rating helpfulness, grading offensiveness — will show lower agreement than an objective one, and that is informative rather than damning. What matters is whether agreement is stable and whether the disagreements cluster somewhere explicable.

Low agreement means bad guidelines

This is the practically useful insight. When agreement is poor, the instinct is to blame annotators and retrain them. Usually the guidelines are ambiguous.

If two careful people reading the same instruction reach different conclusions, the instruction did not determine the answer. The fix is to find the disagreement cluster, write an explicit rule covering it, version the rubric, and re-run. Agreement per label class — rather than one aggregate number — is what makes that cluster visible.

What a delivery report should include

Agreement computed per batch and per label class. Per-annotator accuracy against seeded gold items. The adjudication log showing how disagreements were resolved. The rubric version in force. And the items rejected internally before delivery.

That last one is the tell. A vendor willing to show you what they threw away is a vendor you can calibrate; one reporting only successes is asking for trust you have no basis to give.

Related pages on this site

Answers

Frequently asked questions

Above 0.8 is generally considered strong, 0.67–0.8 workable, below 0.67 grounds for investigation. Adjust expectations by task subjectivity — sentiment or helpfulness ratings legitimately show lower agreement than object detection.

Cohen's kappa handles exactly two annotators on categorical data. Krippendorff's alpha handles any number of annotators, missing data, and nominal, ordinal or interval scales. Alpha is the more practical default for real annotation programmes where not every item is labelled by every annotator.

Look at the disagreements before retraining anyone. Cluster them, and you will usually find a specific case the guidelines never addressed. Write an explicit rule, version the rubric, re-adjudicate the affected items and re-measure. Blaming annotators for ambiguous instructions fixes nothing.

Suspect your labels are noisy?

Send us 500 items and your schema. We will label them free and show you where the disagreements cluster.

Average first response: under 6 business hours. NDAs signed same day.

Free samples Get a quote