Speech & audio data · ready-made dataset
Regional languages speech pack: Saraiki, Balochi, Brahui, Hindko
These are the languages with essentially no public speech corpora. There is no Common Voice of Brahui and no open archive of spoken Saraiki. We collect them where the speakers live, through community networks rather than crowd platforms, because that is the only way to reach them at all.
Datasheet summary
What is in Regional Languages Speech Pack
The least-resourced languages we field, collected in Multan, Quetta and Hazara with community-network recruitment.
| Dataset ID | TDA-SP-004 |
|---|---|
| Modality | Speech |
| Languages / coverage | Saraiki, Balochi, Brahui, Hindko |
| Typical first delivery | 100–300 hrs |
| Free sample | 2–5 hours of audio with transcripts and full metadata, within one business day |
| Provenance | Collected in Pakistan; collection date, location, contributor pseudonym, device profile and consent version on every record |
| Licence | Perpetual, worldwide, sublicensable; derivative-model rights stated explicitly. Exclusivity optional. Licensing detail |
| Delivery | Your S3, GCS or Azure bucket with manifests, checksums, QA report and datasheet |
| Guarantee | 14-day acceptance window; anything outside signed tolerance re-collected free |
Included in every delivery
- Saraiki, Balochi, Brahui and Hindko speech
- Collected in Multan, Quetta and Hazara
- Community-network recruitment rather than crowd platforms
- Transcription by native speakers
Built for
- First-ever ASR coverage for Saraiki, Balochi, Brahui or Hindko
- Extending a multilingual model to languages it currently misroutes
- Language identification and dialect-aware routing for IVR
- Academic and preservation work on under-documented languages
Evaluate before you buy
Start with the free sample, then a pilot
Request 2–5 hours of audio with transcripts and full metadata from Regional Languages Speech Pack — identical in format and quality to the paid corpus, no watermark, no sales call. Pilots are creditable against production and start within five business days of a signed SOW.
Average first response under 6 business hours. NDA signed same day.
- Consent form, licence text and DPA available before you commit
- Datasheet lists composition, QA method and known limitations
- Subsets by language, region or condition on request
- Custom collection to your spec using the same protocol
Answers
Questions about this dataset
Crowd platforms skew heavily toward urban, English-literate users with good bandwidth — a population that barely overlaps with native Saraiki, Balochi, Brahui or Hindko speakers. Community recruitment in Multan, Quetta and Hazara reaches the speakers a crowd platform never will.
A typical first delivery is 100–300 hours. Free sample packs, pilots and larger programmes are all available.
Yes. The pack can be licensed as a whole or as a single-language subset; ask about a subset when you request the sample.
Related datasets
Also in the catalog
Conversational Urdu Speech
Urdu (Karachi, Lahori, standard)
SpeechPakistani Telephony Corpus
Urdu, Punjabi, Sindhi, Pashto
SpeechPakistani-Accented English
English, L1-tagged across 7 languages
See the full catalog · Speech & audio data services · Languages we cover