Text & NLP data · ready-made dataset
RLHF preference pairs
Preference data collected with a translated rubric measures how well raters understood the translation, not how good the response was. These pairs are rated by calibrated native speakers against a rubric written in their own language, with the rationale and agreement metadata delivered alongside every judgment.
Datasheet summary
What is in RLHF Preference Pairs
Side-by-side rankings by calibrated native raters, with rubric scores, rationales and per-rater agreement metadata.
| Dataset ID | TDA-TX-002 |
|---|---|
| Modality | Text |
| Languages / coverage | Urdu, Punjabi, Sindhi, Pashto |
| Typical first delivery | 5k–50k pairs |
| Free sample | 1,000 items with full metadata, within one business day |
| Provenance | Collected in Pakistan; collection date, location, contributor pseudonym, device profile and consent version on every record |
| Licence | Perpetual, worldwide, sublicensable; derivative-model rights stated explicitly. Exclusivity optional. Licensing detail |
| Delivery | Your S3, GCS or Azure bucket with manifests, checksums, QA report and datasheet |
| Guarantee | 14-day acceptance window; anything outside signed tolerance re-collected free |
Included in every delivery
- Side-by-side rankings by calibrated native raters
- Rubric scores and written rationales
- Per-rater agreement metadata
- Urdu, Punjabi, Sindhi and Pashto
Built for
- RLHF and DPO alignment for Pakistani-language models
- Reward-model training with rationale supervision
- Rater-agreement analysis and rubric calibration
- Side-by-side model evaluation in Urdu, Punjabi, Sindhi and Pashto
Evaluate before you buy
Start with the free sample, then a pilot
Request 1,000 items with full metadata from RLHF Preference Pairs — identical in format and quality to the paid corpus, no watermark, no sales call. Pilots are creditable against production and start within five business days of a signed SOW.
Average first response under 6 business hours. NDA signed same day.
- Consent form, licence text and DPA available before you commit
- Datasheet lists composition, QA method and known limitations
- Subsets by language, region or condition on request
- Custom collection to your spec using the same protocol
Answers
Questions about this dataset
Raters are trained on a gold set and monitored for agreement throughout. Per-rater agreement metadata ships with the data so you can weight or filter judgments.
Yes. The same rater panel can run side-by-side evaluations and preference collection on outputs you supply; see the annotation and evaluation service.
The rubric is documented in the datasheet, in the raters' language and in English translation, so your team can audit what each score means.