New: conversational Urdu and Punjabi speech corpora now licensable  See the catalog →

Language guides

Pashto NLP: dialects, script and the data that exists

Pashto spans Pakistan and Afghanistan with two major dialect groups that differ enough to degrade a model tuned on only one. What exists and what a spec must state.

Key takeaways

  • Pashto has roughly 40 million speakers across Pakistan and Afghanistan and very little AI training data.
  • Northern (Yusufzai) and Southern (Kandahari) varieties differ phonologically enough to matter for ASR.
  • The language spans the border with no meaningful discontinuity, so Pakistani-collected data serves Afghan Pashto.
  • Its Perso-Arabic script adds letters not present in Arabic or Urdu, which breaks naive tokenizers and OCR.

Who speaks Pashto and where

Pashto is spoken by roughly 40 million people across Pakistan and Afghanistan. In Pakistan it is concentrated in Khyber Pakhtunkhwa — Peshawar, Mardan, Swat — and the border belt, with large populations in Karachi and other cities through internal migration.

The border is a political line, not a linguistic one. Pashto varieties shade into one another across it, which has a directly practical consequence: data collected in Pakistan serves Afghan Pashto applications well, and this is one of the more reliable transfer relationships in the region.

The dialect split that actually matters

Pashto divides broadly into Northern and Southern groups. Northern — often called Yusufzai, centred on Peshawar — and Southern, often called Kandahari, differ in the realisation of several consonants prominently enough that a model trained on one degrades measurably on the other.

This is not a subtle sociolinguistic distinction to note in a footnote. It is a spec parameter. If your users span both, the corpus must span both, with stated minimum shares.

The script problem

Pashto uses a Perso-Arabic script with additional letters that exist in neither Arabic nor Urdu — retroflex consonants and specific vowel representations among them.

Tokenizers built for Arabic or Urdu mishandle these, either dropping them, mapping them to near-neighbours or fragmenting them badly. OCR models trained on Arabic or Urdu miss them entirely. Any pipeline touching Pashto text needs to be checked explicitly for the full character inventory rather than assumed to work.

What exists today

Pashto appears in some multilingual corpora and academic collections, and volumes are small. Much of what exists is read speech or news text, skewed toward formal register and toward one dialect group without always labelling which.

Unlabelled dialect is a particular problem here: a corpus that does not say whether it is Northern or Southern is difficult to use responsibly, because you cannot tell what you are training toward.

What a Pashto spec should state

Dialect group with minimum shares, explicitly. Region down to district where it matters. Urban and rural split. Gender, which requires care and appropriate recruitment practice in some communities. Device class, since rural connectivity shapes the audio you will actually receive.

And the full character inventory for any text or OCR work, checked against your tokenizer before collection rather than after.

Related pages on this site

Answers

Frequently asked questions

Largely yes. The language spans the border with no meaningful discontinuity, and Northern varieties in particular are shared. Specify the dialect group you need rather than the country, since dialect is the variable that actually affects model performance.

They differ in the realisation of several consonants and in some vocabulary. Northern (Yusufzai) centres on Peshawar; Southern (Kandahari) on Kandahar and the southern border belt. The difference is large enough that an ASR model tuned on one degrades measurably on the other.

Pashto's Perso-Arabic script includes letters absent from both Arabic and Urdu. Tokenizers built for those languages drop, mis-map or badly fragment the additional characters, which degrades representation quality before training begins. Check the full character inventory against your tokenizer explicitly.

Building for Pashto or other Asian-language speakers?

Tell us the language, the dialect group and the task. We will scope it honestly, including the volume ceiling.

Average first response: under 6 business hours. NDAs signed same day.

Free samples Get a quote