Text & NLP data · ready-made dataset
Urdu–English parallel corpus
Most Urdu–English parallel data on the open web is either scraped, machine-translated, or both. This corpus is human-translated and aligned at document and sentence level, with domain glossaries, so a translation model learns real terminology rather than an averaged register.
Datasheet summary
What is in Urdu–English Parallel Corpus
Human-translated document- and sentence-aligned pairs with terminology glossaries across law, medicine, finance, news, government and colloquial registers.
| Dataset ID | TDA-TX-003 |
|---|---|
| Modality | Text |
| Languages / coverage | Urdu↔English, 6 domains |
| Typical first delivery | 100k–1M pairs |
| Free sample | 1,000 items with full metadata, within one business day |
| Provenance | Collected in Pakistan; collection date, location, contributor pseudonym, device profile and consent version on every record |
| Licence | Perpetual, worldwide, sublicensable; derivative-model rights stated explicitly. Exclusivity optional. Licensing detail |
| Delivery | Your S3, GCS or Azure bucket with manifests, checksums, QA report and datasheet |
| Guarantee | 14-day acceptance window; anything outside signed tolerance re-collected free |
Included in every delivery
- Human-translated, document- and sentence-aligned pairs
- Terminology glossaries per domain
- Law, medicine, finance, news, government and colloquial registers
- Both translation directions represented
Built for
- Machine translation training and fine-tuning, Urdu↔English
- Domain adaptation for legal, medical and financial translation
- Cross-lingual retrieval and embedding training
- Terminology-consistent LLM translation evaluation
Evaluate before you buy
Start with the free sample, then a pilot
Request 1,000 items with full metadata from Urdu–English Parallel Corpus — identical in format and quality to the paid corpus, no watermark, no sales call. Pilots are creditable against production and start within five business days of a signed SOW.
Average first response under 6 business hours. NDA signed same day.
- Consent form, licence text and DPA available before you commit
- Datasheet lists composition, QA method and known limitations
- Subsets by language, region or condition on request
- Custom collection to your spec using the same protocol
Answers
Questions about this dataset
Human-translated. Alignment is checked at both document and sentence level before delivery.
Law, medicine, finance, news, government and colloquial text — six registers, each with its own terminology glossary.
A typical first delivery is 100k–1M sentence pairs, and the corpus can be licensed by domain.