Text & NLP data · ready-made dataset
Native Urdu instruction and SFT corpus
Instruction data translated from an English seed set produces a model that sounds foreign — right words, wrong register, examples nobody in Pakistan would give. This corpus is written from scratch by native authors, so the model learns how Urdu speakers actually ask and answer.
Datasheet summary
What is in Native Urdu Instruction & SFT Corpus
Natively authored single- and multi-turn instruction pairs, culturally grounded, with refusal and safety examples.
| Dataset ID | TDA-TX-001 |
|---|---|
| Modality | Text |
| Languages / coverage | Urdu, plus 5 regional languages |
| Typical first delivery | 10k–100k items |
| Free sample | 1,000 items with full metadata, within one business day |
| Provenance | Collected in Pakistan; collection date, location, contributor pseudonym, device profile and consent version on every record |
| Licence | Perpetual, worldwide, sublicensable; derivative-model rights stated explicitly. Exclusivity optional. Licensing detail |
| Delivery | Your S3, GCS or Azure bucket with manifests, checksums, QA report and datasheet |
| Guarantee | 14-day acceptance window; anything outside signed tolerance re-collected free |
Included in every delivery
- Human-authored single- and multi-turn instruction pairs
- Culturally grounded prompts and answers
- Refusal and safety examples
- Urdu, plus five regional languages
Built for
- Supervised fine-tuning of an LLM for Urdu
- Multi-turn chat and assistant behaviour in a South Asian context
- Safety and refusal training on culturally specific cases
- Evaluation sets that a translated benchmark cannot provide
Evaluate before you buy
Start with the free sample, then a pilot
Request 1,000 items with full metadata from Native Urdu Instruction & SFT Corpus — identical in format and quality to the paid corpus, no watermark, no sales call. Pilots are creditable against production and start within five business days of a signed SOW.
Average first response under 6 business hours. NDA signed same day.
- Consent form, licence text and DPA available before you commit
- Datasheet lists composition, QA method and known limitations
- Subsets by language, region or condition on request
- Custom collection to your spec using the same protocol
Answers
Questions about this dataset
No. Every pair is authored natively. Translated instruction data carries English framing into the target language, which is exactly the tone problem this corpus exists to fix.
Instruction fine-tuning changes behaviour usefully from around 10,000 well-authored pairs. A typical first delivery is 10k–100k items.
Urdu plus five regional languages. Ask when requesting a sample if you need a specific language subset.