New: conversational Urdu and Punjabi speech corpora now licensable  See the catalog →

Languages

Every language we collect, in one place

Urdu, Punjabi, Saraiki, Sindhi, Pashto, Balochi, Brahui and more — with the capability tier we can honestly deliver in each, per modality. If your variety is not listed, that usually means a three-week lead time, not a no.

Capability table

What we can deliver, per language and per modality

Deep means a standing panel and production capacity this week. Standard means an active panel with a 2–3 week ramp. On request means bespoke recruitment, scoped honestly before we quote. 18 entries shown.

Language / varietyProvinceWhere spokenSpeakers SpeechTextImage / Video
Urdu (national language)NationwideNationwide~80M (L1+L2)DeepDeepDeep
Punjabi — Majhi (Lahori)PunjabCentral Punjab~40MDeepDeepStandard
Punjabi — Shahpuri & PothwariPunjabN. Punjab, Pothohar~12MStandardStandardOn request
SaraikiPunjabS. Punjab — Multan, Bahawalpur~26MDeepStandardOn request
Sindhi — Vicholi (standard)SindhCentral Sindh~32MDeepDeepStandard
Sindhi — Lari & ThariSindhS. Sindh, Tharparkar~6MStandardOn requestOn request
Pashto — Northern (Yusufzai)Khyber PakhtunkhwaPeshawar, Mardan, Swat~28MDeepStandardStandard
Pashto — Southern (Kandahari)Khyber PakhtunkhwaWaziristan, border belt~12MStandardOn requestOn request
Balochi — Rakhshani & MakraniBalochistanQuetta, Turbat, Gwadar~8MStandardStandardOn request
BrahuiBalochistanKalat, Mastung~2.8MStandardOn requestOn request
HindkoKhyber PakhtunkhwaHazara, Abbottabad, Peshawar city~5MStandardOn requestOn request
Kashmiri & Pahari-PothwariAzad Jammu & KashmirAJK, Muzaffarabad~4MStandardOn requestOn request
ShinaGilgit-BaltistanGilgit, Chilas~1MOn requestOn requestOn request
BaltiGilgit-BaltistanSkardu, Baltistan~0.4MOn requestOn requestOn request
KhowarGilgit-BaltistanChitral~0.3MOn requestOn requestOn request
BurushaskiGilgit-BaltistanHunza, Nagar, Yasin~0.1MOn requestOn requestOn request
Pakistani-accented EnglishNationwideNationwide, CEFR-graded~110M (L2)DeepDeepStandard
Urdu–English code-switchingNationwideUrban nationwideDeepDeepOn request

Speaker figures are approximate, including second-language speakers, drawn from the 2023 census and public sources.

Urdu

اردو

roughly 80 million first- and second-language speakers

Collected in: Nationwide, with the deepest urban pools in Karachi, Lahore and Islamabad

Request Urdu data

Urdu is Pakistan's national language and the lingua franca between provinces. It is written in the Nastaliq style of the Perso-Arabic script, which is why off-the-shelf OCR trained on Naskh Arabic fails on Urdu documents. Spoken Urdu is also largely mutually intelligible with spoken Hindi, so Urdu speech data carries measurable transfer value into Hindi ASR.

ASR & TTSLLM fine-tuning TranslationOCR Annotation

Pashto

پښتو

roughly 40 million speakers across Pakistan and Afghanistan

Collected in: Khyber Pakhtunkhwa — Peshawar, Mardan, Swat — and the border belt

Request Pashto data

Pashto splits into Northern (Yusufzai) and Southern (Kandahari) varieties that differ enough in phonology to degrade a model tuned on only one. It spans the Afghan border with no meaningful linguistic discontinuity, so Pashto data collected in Pakistan transfers directly to Afghan Pashto and closely to Dari.

ASR & TTSLLM fine-tuning TranslationOCR Annotation

Punjabi

پنجابی

roughly 90 million speakers in Pakistan alone

Collected in: Punjab — Lahore, Faisalabad, Gujranwala, Sialkot and rural districts

Request Punjabi data

Punjabi is the most spoken language in Pakistan and one of the most spoken in the world, yet among the least resourced in AI. In Pakistan it is written in Shahmukhi; across the Indian border in Gurmukhi. The scripts differ but the speech does not, so Punjabi audio collected in Pakistan transfers almost intact to Indian Punjabi speech tasks.

ASR & TTSLLM fine-tuning TranslationOCR Annotation

Sindhi

سنڌي

roughly 38 million speakers

Collected in: Sindh — Karachi, Hyderabad, Sukkur, Larkana and rural Sindh

Request Sindhi data

Sindhi uses an extended Perso-Arabic alphabet with 52 letters, more than any other Arabic-script language, which breaks tokenizers and OCR built for Arabic or Urdu. The Vicholi variety is the written standard; Lari and Thari differ enough in the south to warrant separate quotas.

ASR & TTSLLM fine-tuning TranslationOCR Annotation

Saraiki, Balochi

سرائیکی · بلوچی · براہوئی

roughly 37 million speakers combined

Collected in: Southern Punjab (Multan, Bahawalpur) and Balochistan (Quetta, Turbat, Kalat)

Request Saraiki, Balochi data

These are the genuinely low-resource languages, where almost no usable training data exists in any public corpus. Brahui is especially unusual — a Dravidian language surrounded by Indo-Iranian ones, with no meaningful transfer from any neighbouring language. If your model needs to serve these speakers, the data has to be collected from scratch, and very few organisations can field it.

ASR & TTSLLM fine-tuning TranslationOCR Annotation

Every modality

What we deliver in any of these languages

Data typeWhat it containsWhat it is used for
Conversational speech (ASR)Spontaneous two-party dialogue, verbatim transcripts, diarizationSpeech recognition, voice agents, call analytics
Scripted speech (TTS)Phonetically balanced prompts, studio capture, single-speakerText-to-speech, voice cloning, pronunciation models
Telephony audio8 kHz narrowband, channel-separated, GSM and VoIP codecsIVR, call-centre automation, fraud detection
Native text corporaOriginal long-form writing across registers and domainsPretraining, continued pretraining, language modelling
Instruction / SFT pairsNatively authored prompts and responses, multi-turnSupervised fine-tuning, assistant alignment
Preference / RLHF dataSide-by-side rankings by calibrated native ratersReward modelling, RLHF, DPO
Translation pairsHuman translation to and from English, document-alignedMachine translation, cross-lingual transfer
OCR & handwritingPrinted and handwritten documents with reading-order labelsDocument AI, information extraction
Safety & red-teamCulturally specific toxicity, slurs, jailbreaks, severity-gradedSafety classifiers, guardrails, evaluation

Answers

Language data questions

Urdu, Punjabi (Majhi, Shahpuri and Pothwari), Saraiki, Sindhi (Vicholi, Lari and Thari), Pashto (Northern and Southern), Balochi, Brahui, Hindko, Kashmiri, Shina, Balti, Khowar and Burushaski, plus Pakistani-accented English and Urdu–English code-switched speech. The table above states our current capability tier per language and per modality.

Directly from The Dataa. We collect at source in Pakistan to a written specification, with per-record informed consent and a perpetual, worldwide, sublicensable licence that names machine-learning training explicitly. Free unwatermarked samples arrive within one business day and no sales call is required to get them.

Rough guidance from projects we have run: fine-tuning a multilingual ASR model shows meaningful gains from 200–500 hours and keeps improving to a few thousand. A single-speaker TTS voice needs 20–60 studio hours. Instruction fine-tuning changes behaviour usefully from around 10,000 well-authored pairs. Training from scratch is a different order of magnitude and rarely the right call for a low-resource language.

Some, and you should exhaust it before paying anyone. Common Voice, FLEURS and OpenSLR carry Urdu, Punjabi, Pashto and Sindhi material. The limits are consistent: small volumes, narrow speaker demographics, read rather than spontaneous speech, and licences that are often unclear for commercial model training. Commissioned collection is what you buy when the open data has run out or its licence will not survive review.

Usually yes. The table shows where we already have standing panels; it is not the limit of what we can field. A new language typically needs 3–5 weeks of extra lead time to recruit, screen and calibrate a local team. Tell us the language and region and we will come back within two business days with a feasibility answer, a timeline and an honest view of the volume ceiling.

Yes, and we usually insist on it. “Punjabi” alone spans Majhi, Shahpuri, Pothwari and Jhangvi varieties that differ enough to break an ASR model tuned on Lahori newsreader speech, and Saraiki is a separate language again. Our specs name dialects and regions explicitly, and we recruit to district level where you need that resolution.

Name the language. We will name the timeline.

Two business days to a feasibility answer, a realistic volume ceiling and a scoped approach.

Average first response: under 6 business hours. NDAs signed same day.

Free samples Get a quote