New: conversational Urdu and Punjabi speech corpora now licensable  See the catalog →

Text & NLP data · ready-made dataset

Native Urdu instruction and SFT corpus

Instruction data translated from an English seed set produces a model that sounds foreign — right words, wrong register, examples nobody in Pakistan would give. This corpus is written from scratch by native authors, so the model learns how Urdu speakers actually ask and answer.

Free unwatermarked samplePer-record consentTraining-rights licence14-day acceptance

Datasheet summary

What is in Native Urdu Instruction & SFT Corpus

Natively authored single- and multi-turn instruction pairs, culturally grounded, with refusal and safety examples.

Dataset IDTDA-TX-001
ModalityText
Languages / coverageUrdu, plus 5 regional languages
Typical first delivery10k–100k items
Free sample1,000 items with full metadata, within one business day
ProvenanceCollected in Pakistan; collection date, location, contributor pseudonym, device profile and consent version on every record
LicencePerpetual, worldwide, sublicensable; derivative-model rights stated explicitly. Exclusivity optional. Licensing detail
DeliveryYour S3, GCS or Azure bucket with manifests, checksums, QA report and datasheet
Guarantee14-day acceptance window; anything outside signed tolerance re-collected free

Included in every delivery

  • Human-authored single- and multi-turn instruction pairs
  • Culturally grounded prompts and answers
  • Refusal and safety examples
  • Urdu, plus five regional languages

Built for

  • Supervised fine-tuning of an LLM for Urdu
  • Multi-turn chat and assistant behaviour in a South Asian context
  • Safety and refusal training on culturally specific cases
  • Evaluation sets that a translated benchmark cannot provide

Evaluate before you buy

Start with the free sample, then a pilot

Request 1,000 items with full metadata from Native Urdu Instruction & SFT Corpus — identical in format and quality to the paid corpus, no watermark, no sales call. Pilots are creditable against production and start within five business days of a signed SOW.

Average first response under 6 business hours. NDA signed same day.

  • Consent form, licence text and DPA available before you commit
  • Datasheet lists composition, QA method and known limitations
  • Subsets by language, region or condition on request
  • Custom collection to your spec using the same protocol

Answers

Questions about this dataset

No. Every pair is authored natively. Translated instruction data carries English framing into the target language, which is exactly the tone problem this corpus exists to fix.

Instruction fine-tuning changes behaviour usefully from around 10,000 well-authored pairs. A typical first delivery is 10k–100k items.

Urdu plus five regional languages. Ask when requesting a sample if you need a specific language subset.

Free samples Get a quote