New: conversational Urdu and Punjabi speech corpora now licensable  See the catalog →

Language guides

Punjabi NLP: 90 million speakers, almost no data

Punjabi is among the world's most spoken languages and among the least resourced in AI. What exists, what does not, and how Shahmukhi and Gurmukhi complicate everything.

Key takeaways

  • Punjabi has roughly 90 million speakers in Pakistan alone and is severely under-resourced in AI.
  • Two scripts: Shahmukhi in Pakistan, Gurmukhi in India. The speech is shared; the writing is not.
  • Dialects — Majhi, Shahpuri, Pothwari, Jhangvi — differ enough to affect ASR performance materially.
  • Speech data collected on either side of the border transfers well; text and OCR data does not.

The scale-versus-resource gap

Punjabi is spoken by roughly 90 million people in Pakistan and tens of millions more in India and the diaspora, placing it among the most spoken languages on earth. Its representation in AI training corpora is a rounding error.

The reason is structural rather than mysterious. In Pakistan, Urdu is the language of education, administration and most publishing, so Punjabi is enormously spoken and comparatively little written. Digital corpus size tracks written output, not speakers.

Two scripts, one spoken language

In Pakistan, Punjabi is written in Shahmukhi, a Perso-Arabic script closely related to Urdu's Nastaliq. In India, it is written in Gurmukhi, a Brahmic script unrelated to it.

The consequence for AI work is clean and important. Speech data transfers almost intact across the border — the spoken language is essentially shared. Text, OCR and anything script-dependent does not transfer at all. Teams routinely conflate the two and buy the wrong thing.

The dialects that matter

Majhi, centred on Lahore and Amritsar, is the prestige variety and the basis of most written standard. Shahpuri and Pothwari in northern Punjab and the Pothohar plateau differ noticeably. Jhangvi occupies the centre.

And Saraiki, spoken across southern Punjab, is generally classified as a separate language rather than a Punjabi dialect — a distinction that matters when writing a spec, because a Saraiki speaker is not a Punjabi data point.

What exists today

Some Punjabi material appears in Common Voice, FLEURS and various academic corpora, weighted toward Gurmukhi and Indian Punjabi. Shahmukhi Punjabi from Pakistan is substantially thinner.

The consistent limits: modest volume, read rather than spontaneous speech, self-selected urban contributors, and licences that were rarely drafted with commercial model training in mind. Useful for prototyping; usually insufficient for production.

What a Punjabi collection spec should specify

Script, unambiguously — Shahmukhi or Gurmukhi, and if both, how they are labelled. Dialect shares with minimum percentages. Urban and rural split, since rural Punjabi differs meaningfully from urban. Age bands, because younger urban speakers code-switch into Urdu and English far more.

And an explicit decision on Saraiki: included as a separate language with its own quota, or excluded. Leaving it ambiguous produces a corpus nobody can interpret afterwards.

Related pages on this site

Answers

Frequently asked questions

Yes, despite roughly 90 million speakers in Pakistan alone. Low-resource refers to available machine-readable data, not speaker population. Punjabi is widely spoken and comparatively little written, particularly in Pakistan where Urdu dominates education and publishing.

They are two writing systems for the same spoken language. Shahmukhi is Perso-Arabic and used in Pakistan; Gurmukhi is Brahmic and used in India. Speech data transfers between them almost intact. Text and OCR data does not transfer at all.

For speech, largely yes — the spoken language is essentially shared, with some lexical borrowing differences. For text and OCR, no, because of the script difference. Check which script any corpus actually uses before assuming it fits your task.

Building for Punjabi or other Asian-language speakers?

Tell us the script, the dialects and the task. We will come back with a spec and a realistic timeline.

Average first response: under 6 business hours. NDAs signed same day.

Free samples Get a quote