New: conversational Urdu and Punjabi speech corpora now licensable  See the catalog →

Language coverage

Pashto training datasets

Pashto is spoken across a border by tens of millions of people, in dialect groups different enough to break a model tuned on only one. What that means for collection, labelling and any spec worth signing.

Key takeaways

  • Pashto has two major dialect groups. Unlabelled data mixing them is worth less than either group collected cleanly.
  • Orthography varies by region and education. Without an enforced convention, one word becomes several tokens.
  • Most open Pashto audio is read speech from a narrow demographic, which is the opposite of production conditions.
  • Collecting from women contributors requires deliberate design and costs more. A corpus without it fails half the user base.
  • Cross-border sourcing raises real consent and provenance questions. Ask how they were answered, per record.

The dialect problem comes first

Every other decision about Pashto data is downstream of one question: which Pashto. The language divides into two major groups whose pronunciation differs systematically, and the difference is large enough that an acoustic model trained on one degrades measurably on the other. Vocabulary and borrowed terminology diverge as well, more so across the Pakistan and Afghanistan border than within either country.

The failure mode this produces is unusually expensive to diagnose. A model tuned on one group and deployed against the other does not crash. It returns fluent, confident transcripts that are wrong in scattered places, and the error concentrates in exactly the words that carry meaning. Teams often chase this for months as a decoding or language-model problem when it was a sampling problem from the start.

The fix is unglamorous. Decide which group you serve, or collect both in a stated proportion, and label the group on every single recording so that you can evaluate them separately.

Orthography, and why it quietly ruins corpora

Pashto is written in a Perso-Arabic script with letters that exist in no other language, and its spelling conventions are not uniformly settled. Regional habit and level of schooling both influence how a word is written. Two transcribers working in good faith on the same audio will produce different strings.

Left unmanaged, this fragments the vocabulary. The model sees several spellings of a common word, treats them as unrelated tokens, and learns each one weakly instead of one of them well. Nothing in the delivery looks wrong; the corpus simply underperforms its size.

The remedy is a written orthographic convention issued to transcribers before work starts, and enforced by a second reviewer who checks conformance rather than only meaning. This is the single highest-leverage decision on a Pashto project and it costs nothing but attention.

Why open Pashto data underperforms

What you can find openlyUseful forWhy production still fails
Read speech from broadcast or scripted sourcesPhoneme coverage, an initial acoustic priorNo spontaneity, narrow speaker pool, no telephone band
News and web textLanguage modelling, vocabularyFormal register only, dialect unlabelled, inconsistent spelling
Multilingual checkpoints that list PashtoA starting point for fine-tuningThin Pashto share; fluency masks systematic dialect error
Community and volunteer recordingsGoodwill, some diversityConsent provenance rarely evidenced per record; uneven quality

The common thread is that open Pashto is skewed towards the formal, the male, the urban and the read. Production systems meet the opposite of all four.

Collecting Pashto responsibly

Pashto collection raises questions that a generic vendor questionnaire will not surface, and they are worth asking directly.

  • Where were the contributors, and under what circumstances? Cross-border and displaced populations can be recruited cheaply and cannot always give consent freely. Ask where recruitment happened and what the alternative to participating was.
  • How was consent explained? Literacy rates vary widely across the Pashto-speaking population. A signed English form is not evidence of informed consent, and a vendor that offers one as proof is telling you something.
  • Were women contributors included, and how? This requires female recruiters and session leads, venues a household will accept, and more time per session. It costs more. It is also the difference between a system that serves the population and one that serves half of it.
  • What were people paid, and against what benchmark? Ask for the rate and the local reference point, not an assurance that pay was fair.
  • Can consent be withdrawn after delivery? If yes, ask what happens to your copy and your trained weights. If the answer is vague, the licence is vague.

We publish our answers to all five rather than supplying them on request. Consent and fair pay sets out the standard, and trust and compliance covers the documentation delivered with every order.

What a Pashto specification has to state

  • Dialect group, and the proportion of each if you want both, with the label delivered per record.
  • Country of collection, stated explicitly rather than implied by the language name.
  • Orthographic convention, issued in writing with worked examples of contested spellings.
  • Speaker quotas on gender, age band, region and urban or rural origin, with a tolerance on each.
  • Acoustic conditions, including narrowband capture if the deployment is telephony.
  • Code-switching policy for Urdu, Dari and English fragments, which occur constantly in real speech.
  • Evaluation split held out per dialect group, so that a gain on one cannot hide a loss on the other.
  • Acceptance test and threshold, with the consequence of a miss written into the contract.

Pashto sits alongside Sindhi, Saraiki and Balochi in our regional languages speech pack, and can be commissioned on its own to a dedicated spec.

Related pages on this site

FAQ

Questions buyers ask about Pashto data

The one your users speak. If you do not know, collect both major groups in a stated proportion and label every recording with its dialect group. A model tuned on one group and deployed against the other loses accuracy in exactly the way that is hardest to debug, because the transcripts look plausible.

Not for training purposes. Pronunciation, some vocabulary and a good deal of borrowed terminology differ, and the media diet that shapes formal register differs too. Both are legitimate Pashto. A corpus that mixes them without labels is less useful than either one collected cleanly.

By writing the orthographic convention into the specification and enforcing it in the second QA pass rather than assuming transcribers share one. Pashto orthography varies by region and education, so an unmanaged corpus will contain several spellings of common words and teach the model that they are unrelated tokens.

Yes, and it requires deliberate design: female recruiters, female session leads, appropriate venues and consent processes that a household will accept. It costs more per hour than collecting from men, and a corpus without it produces a system that fails half its users.

Enough for a serious fine-tune, collected to a stated quota, is achievable on a normal project timeline. What is not realistic is expecting an off-the-shelf Pashto corpus to match your production conditions, because the open material is thin, dialect-unlabelled and mostly read speech.

Pashto samples, labelled by dialect group.

Back within one business day, with the consent record and the orthographic convention attached.

Free samples Get a quote