New: conversational Urdu and Punjabi speech corpora now licensable  See the catalog →

Research

RLHF outside English: why translated rubrics fail

Collecting preference data outside English needs calibrated native raters and a rubric written in their language. Why translated rubrics produce noise, and how to fix it.

Key takeaways

  • Preference data quality depends on rater calibration, and calibration does not survive translation of the rubric.
  • Rater agreement should be measured and reported. Unmeasured preference data is noise you paid for.
  • “Helpful” and “polite” are culturally loaded and need locally-grounded definitions.
  • Safety-adjacent rating needs a separate, better-paid pool with welfare protections and unconditional opt-out.

Why preference data is harder than it looks

Reward modelling and DPO both depend on humans consistently ranking outputs. Consistency is the operative word: if raters disagree at random, the reward model learns the average of noise and the resulting policy drifts in ways nobody intended.

In English, with a mature rubric and an experienced rater pool, agreement is achievable. Outside English, with a rubric translated last week and raters recruited last month, it usually is not — and the resulting dataset looks superficially fine while being close to worthless.

The translated rubric problem

A rating rubric is not a document; it is a shared understanding. Translating the words does not transfer the understanding, because the terms carry culturally specific content.

“Helpful” means something different where directness is valued versus where it reads as rude. “Polite” in Urdu involves honorific register choices with no English equivalent. “Harmful” depends on local legal and social context. Raters given a translated rubric interpret these terms individually, and their individual interpretations diverge — which shows up as low agreement.

What calibration actually requires

A rubric authored in the target language by someone who knows both the domain and the culture, rather than translated. A mandatory qualification round against pre-adjudicated gold items before a rater touches live work. Ongoing gold-question injection at a set rate so drift is caught early. And a senior adjudicator whose rulings feed back into a versioned rubric.

Rater agreement should then be measured and reported per batch — Cohen's kappa or Krippendorff's alpha depending on the schema — so you can see which categories are genuinely ambiguous rather than assuming uniform quality.

What to ask a vendor for

Per-rater accuracy distribution, not an average. Inter-rater agreement by label class, so you can see where the rubric is unclear. The adjudication log. The rubric version that governed each batch. And the items rejected internally before delivery.

A vendor reporting only a single headline quality number is a vendor you cannot calibrate against, which means you cannot tell good batches from bad ones.

Safety work needs different treatment

Red-teaming, toxicity rating and harm-taxonomy work expose raters to genuinely distressing material. Treating that pool identically to a general annotation pool is both an ethical failure and a data quality problem, because distressed and fatigued raters produce inconsistent judgements.

The protections that matter: enhanced screening at recruitment, briefing and debriefing, capped exposure hours, paid rest built into the schedule, access to support, and genuinely unconditional opt-out at any point without consequence.

Related pages on this site

Answers

Frequently asked questions

You can translate the words; the shared understanding does not come with them. Terms like helpful, polite and harmful are culturally loaded, so translated rubrics produce divergent individual interpretations and low inter-rater agreement. A rubric authored in the target language by someone who knows the culture works considerably better.

Through inter-rater agreement — Cohen's kappa or Krippendorff's alpha — computed per batch and ideally per label class, plus per-rater accuracy against invisibly seeded gold questions. Both should be reported to you rather than summarised into a single number.

It varies with task breadth and base model strength, but tens of thousands of pairs is a common operating range for a focused domain. Rater consistency affects the outcome more than volume does: noisy preferences at scale produce a reward model that confidently learns the noise.

Send us 500 items and your rubric.

We will rate them, return the agreement numbers, and show you the disagreements. If the numbers are not better than what you have, we have not earned the work.

Average first response: under 6 business hours. NDAs signed same day.

Free samples Get a quote