Languages
Every language we collect, in one place
Urdu, Punjabi, Saraiki, Sindhi, Pashto, Balochi, Brahui and more — with the capability tier we can honestly deliver in each, per modality. If your variety is not listed, that usually means a three-week lead time, not a no.
Capability table
What we can deliver, per language and per modality
Deep means a standing panel and production capacity this week. Standard means an active panel with a 2–3 week ramp. On request means bespoke recruitment, scoped honestly before we quote. 18 entries shown.
| Language / variety | Province | Where spoken | Speakers | Speech | Text | Image / Video |
|---|---|---|---|---|---|---|
| Urdu (national language) | Nationwide | Nationwide | ~80M (L1+L2) | Deep | Deep | Deep |
| Punjabi — Majhi (Lahori) | Punjab | Central Punjab | ~40M | Deep | Deep | Standard |
| Punjabi — Shahpuri & Pothwari | Punjab | N. Punjab, Pothohar | ~12M | Standard | Standard | On request |
| Saraiki | Punjab | S. Punjab — Multan, Bahawalpur | ~26M | Deep | Standard | On request |
| Sindhi — Vicholi (standard) | Sindh | Central Sindh | ~32M | Deep | Deep | Standard |
| Sindhi — Lari & Thari | Sindh | S. Sindh, Tharparkar | ~6M | Standard | On request | On request |
| Pashto — Northern (Yusufzai) | Khyber Pakhtunkhwa | Peshawar, Mardan, Swat | ~28M | Deep | Standard | Standard |
| Pashto — Southern (Kandahari) | Khyber Pakhtunkhwa | Waziristan, border belt | ~12M | Standard | On request | On request |
| Balochi — Rakhshani & Makrani | Balochistan | Quetta, Turbat, Gwadar | ~8M | Standard | Standard | On request |
| Brahui | Balochistan | Kalat, Mastung | ~2.8M | Standard | On request | On request |
| Hindko | Khyber Pakhtunkhwa | Hazara, Abbottabad, Peshawar city | ~5M | Standard | On request | On request |
| Kashmiri & Pahari-Pothwari | Azad Jammu & Kashmir | AJK, Muzaffarabad | ~4M | Standard | On request | On request |
| Shina | Gilgit-Baltistan | Gilgit, Chilas | ~1M | On request | On request | On request |
| Balti | Gilgit-Baltistan | Skardu, Baltistan | ~0.4M | On request | On request | On request |
| Khowar | Gilgit-Baltistan | Chitral | ~0.3M | On request | On request | On request |
| Burushaski | Gilgit-Baltistan | Hunza, Nagar, Yasin | ~0.1M | On request | On request | On request |
| Pakistani-accented English | Nationwide | Nationwide, CEFR-graded | ~110M (L2) | Deep | Deep | Standard |
| Urdu–English code-switching | Nationwide | Urban nationwide | — | Deep | Deep | On request |
Speaker figures are approximate, including second-language speakers, drawn from the 2023 census and public sources.
Urdu
اردو
roughly 80 million first- and second-language speakers
Collected in: Nationwide, with the deepest urban pools in Karachi, Lahore and Islamabad
Urdu is Pakistan's national language and the lingua franca between provinces. It is written in the Nastaliq style of the Perso-Arabic script, which is why off-the-shelf OCR trained on Naskh Arabic fails on Urdu documents. Spoken Urdu is also largely mutually intelligible with spoken Hindi, so Urdu speech data carries measurable transfer value into Hindi ASR.
Pashto
پښتو
roughly 40 million speakers across Pakistan and Afghanistan
Collected in: Khyber Pakhtunkhwa — Peshawar, Mardan, Swat — and the border belt
Pashto splits into Northern (Yusufzai) and Southern (Kandahari) varieties that differ enough in phonology to degrade a model tuned on only one. It spans the Afghan border with no meaningful linguistic discontinuity, so Pashto data collected in Pakistan transfers directly to Afghan Pashto and closely to Dari.
Punjabi
پنجابی
roughly 90 million speakers in Pakistan alone
Collected in: Punjab — Lahore, Faisalabad, Gujranwala, Sialkot and rural districts
Punjabi is the most spoken language in Pakistan and one of the most spoken in the world, yet among the least resourced in AI. In Pakistan it is written in Shahmukhi; across the Indian border in Gurmukhi. The scripts differ but the speech does not, so Punjabi audio collected in Pakistan transfers almost intact to Indian Punjabi speech tasks.
Sindhi
سنڌي
roughly 38 million speakers
Collected in: Sindh — Karachi, Hyderabad, Sukkur, Larkana and rural Sindh
Sindhi uses an extended Perso-Arabic alphabet with 52 letters, more than any other Arabic-script language, which breaks tokenizers and OCR built for Arabic or Urdu. The Vicholi variety is the written standard; Lari and Thari differ enough in the south to warrant separate quotas.
Saraiki, Balochi
سرائیکی · بلوچی · براہوئی
roughly 37 million speakers combined
Collected in: Southern Punjab (Multan, Bahawalpur) and Balochistan (Quetta, Turbat, Kalat)
These are the genuinely low-resource languages, where almost no usable training data exists in any public corpus. Brahui is especially unusual — a Dravidian language surrounded by Indo-Iranian ones, with no meaningful transfer from any neighbouring language. If your model needs to serve these speakers, the data has to be collected from scratch, and very few organisations can field it.
Every modality
What we deliver in any of these languages
| Data type | What it contains | What it is used for |
|---|---|---|
| Conversational speech (ASR) | Spontaneous two-party dialogue, verbatim transcripts, diarization | Speech recognition, voice agents, call analytics |
| Scripted speech (TTS) | Phonetically balanced prompts, studio capture, single-speaker | Text-to-speech, voice cloning, pronunciation models |
| Telephony audio | 8 kHz narrowband, channel-separated, GSM and VoIP codecs | IVR, call-centre automation, fraud detection |
| Native text corpora | Original long-form writing across registers and domains | Pretraining, continued pretraining, language modelling |
| Instruction / SFT pairs | Natively authored prompts and responses, multi-turn | Supervised fine-tuning, assistant alignment |
| Preference / RLHF data | Side-by-side rankings by calibrated native raters | Reward modelling, RLHF, DPO |
| Translation pairs | Human translation to and from English, document-aligned | Machine translation, cross-lingual transfer |
| OCR & handwriting | Printed and handwritten documents with reading-order labels | Document AI, information extraction |
| Safety & red-team | Culturally specific toxicity, slurs, jailbreaks, severity-graded | Safety classifiers, guardrails, evaluation |
Answers
Language data questions
Urdu, Punjabi (Majhi, Shahpuri and Pothwari), Saraiki, Sindhi (Vicholi, Lari and Thari), Pashto (Northern and Southern), Balochi, Brahui, Hindko, Kashmiri, Shina, Balti, Khowar and Burushaski, plus Pakistani-accented English and Urdu–English code-switched speech. The table above states our current capability tier per language and per modality.
Directly from The Dataa. We collect at source in Pakistan to a written specification, with per-record informed consent and a perpetual, worldwide, sublicensable licence that names machine-learning training explicitly. Free unwatermarked samples arrive within one business day and no sales call is required to get them.
Rough guidance from projects we have run: fine-tuning a multilingual ASR model shows meaningful gains from 200–500 hours and keeps improving to a few thousand. A single-speaker TTS voice needs 20–60 studio hours. Instruction fine-tuning changes behaviour usefully from around 10,000 well-authored pairs. Training from scratch is a different order of magnitude and rarely the right call for a low-resource language.
Some, and you should exhaust it before paying anyone. Common Voice, FLEURS and OpenSLR carry Urdu, Punjabi, Pashto and Sindhi material. The limits are consistent: small volumes, narrow speaker demographics, read rather than spontaneous speech, and licences that are often unclear for commercial model training. Commissioned collection is what you buy when the open data has run out or its licence will not survive review.
Usually yes. The table shows where we already have standing panels; it is not the limit of what we can field. A new language typically needs 3–5 weeks of extra lead time to recruit, screen and calibrate a local team. Tell us the language and region and we will come back within two business days with a feasibility answer, a timeline and an honest view of the volume ceiling.
Yes, and we usually insist on it. “Punjabi” alone spans Majhi, Shahpuri, Pothwari and Jhangvi varieties that differ enough to break an ASR model tuned on Lahori newsreader speech, and Saraiki is a separate language again. Our specs name dialects and regions explicitly, and we recruit to district level where you need that resolution.
Name the language. We will name the timeline.
Two business days to a feasibility answer, a realistic volume ceiling and a scoped approach.
Average first response: under 6 business hours. NDAs signed same day.