New: conversational Urdu and Punjabi speech corpora now licensable  See the catalog →

Text & NLP data · ready-made dataset

RLHF preference pairs

Preference data collected with a translated rubric measures how well raters understood the translation, not how good the response was. These pairs are rated by calibrated native speakers against a rubric written in their own language, with the rationale and agreement metadata delivered alongside every judgment.

Free unwatermarked samplePer-record consentTraining-rights licence14-day acceptance

Datasheet summary

What is in RLHF Preference Pairs

Side-by-side rankings by calibrated native raters, with rubric scores, rationales and per-rater agreement metadata.

Dataset IDTDA-TX-002
ModalityText
Languages / coverageUrdu, Punjabi, Sindhi, Pashto
Typical first delivery5k–50k pairs
Free sample1,000 items with full metadata, within one business day
ProvenanceCollected in Pakistan; collection date, location, contributor pseudonym, device profile and consent version on every record
LicencePerpetual, worldwide, sublicensable; derivative-model rights stated explicitly. Exclusivity optional. Licensing detail
DeliveryYour S3, GCS or Azure bucket with manifests, checksums, QA report and datasheet
Guarantee14-day acceptance window; anything outside signed tolerance re-collected free

Included in every delivery

  • Side-by-side rankings by calibrated native raters
  • Rubric scores and written rationales
  • Per-rater agreement metadata
  • Urdu, Punjabi, Sindhi and Pashto

Built for

  • RLHF and DPO alignment for Pakistani-language models
  • Reward-model training with rationale supervision
  • Rater-agreement analysis and rubric calibration
  • Side-by-side model evaluation in Urdu, Punjabi, Sindhi and Pashto

Evaluate before you buy

Start with the free sample, then a pilot

Request 1,000 items with full metadata from RLHF Preference Pairs — identical in format and quality to the paid corpus, no watermark, no sales call. Pilots are creditable against production and start within five business days of a signed SOW.

Average first response under 6 business hours. NDA signed same day.

  • Consent form, licence text and DPA available before you commit
  • Datasheet lists composition, QA method and known limitations
  • Subsets by language, region or condition on request
  • Custom collection to your spec using the same protocol

Answers

Questions about this dataset

Raters are trained on a gold set and monitored for agreement throughout. Per-rater agreement metadata ships with the data so you can weight or filter judgments.

Yes. The same rater panel can run side-by-side evaluations and preference collection on outputs you supply; see the annotation and evaluation service.

The rubric is documented in the datasheet, in the raters' language and in English translation, so your team can audit what each score means.

Free samples Get a quote