New: conversational Urdu and Punjabi speech corpora now licensable  See the catalog →

Text & NLP data · ready-made dataset

Urdu–English parallel corpus

Most Urdu–English parallel data on the open web is either scraped, machine-translated, or both. This corpus is human-translated and aligned at document and sentence level, with domain glossaries, so a translation model learns real terminology rather than an averaged register.

Free unwatermarked samplePer-record consentTraining-rights licence14-day acceptance

Datasheet summary

What is in Urdu–English Parallel Corpus

Human-translated document- and sentence-aligned pairs with terminology glossaries across law, medicine, finance, news, government and colloquial registers.

Dataset IDTDA-TX-003
ModalityText
Languages / coverageUrdu↔English, 6 domains
Typical first delivery100k–1M pairs
Free sample1,000 items with full metadata, within one business day
ProvenanceCollected in Pakistan; collection date, location, contributor pseudonym, device profile and consent version on every record
LicencePerpetual, worldwide, sublicensable; derivative-model rights stated explicitly. Exclusivity optional. Licensing detail
DeliveryYour S3, GCS or Azure bucket with manifests, checksums, QA report and datasheet
Guarantee14-day acceptance window; anything outside signed tolerance re-collected free

Included in every delivery

  • Human-translated, document- and sentence-aligned pairs
  • Terminology glossaries per domain
  • Law, medicine, finance, news, government and colloquial registers
  • Both translation directions represented

Built for

  • Machine translation training and fine-tuning, Urdu↔English
  • Domain adaptation for legal, medical and financial translation
  • Cross-lingual retrieval and embedding training
  • Terminology-consistent LLM translation evaluation

Evaluate before you buy

Start with the free sample, then a pilot

Request 1,000 items with full metadata from Urdu–English Parallel Corpus — identical in format and quality to the paid corpus, no watermark, no sales call. Pilots are creditable against production and start within five business days of a signed SOW.

Average first response under 6 business hours. NDA signed same day.

  • Consent form, licence text and DPA available before you commit
  • Datasheet lists composition, QA method and known limitations
  • Subsets by language, region or condition on request
  • Custom collection to your spec using the same protocol

Answers

Questions about this dataset

Human-translated. Alignment is checked at both document and sentence level before delivery.

Law, medicine, finance, news, government and colloquial text — six registers, each with its own terminology glossary.

A typical first delivery is 100k–1M sentence pairs, and the corpus can be licensed by domain.

Free samples Get a quote