New: conversational Urdu and Punjabi speech corpora now licensable  See the catalog →

Speech & audio data · ready-made dataset

Regional languages speech pack: Saraiki, Balochi, Brahui, Hindko

These are the languages with essentially no public speech corpora. There is no Common Voice of Brahui and no open archive of spoken Saraiki. We collect them where the speakers live, through community networks rather than crowd platforms, because that is the only way to reach them at all.

Free unwatermarked samplePer-record consentTraining-rights licence14-day acceptance

Datasheet summary

What is in Regional Languages Speech Pack

The least-resourced languages we field, collected in Multan, Quetta and Hazara with community-network recruitment.

Dataset IDTDA-SP-004
ModalitySpeech
Languages / coverageSaraiki, Balochi, Brahui, Hindko
Typical first delivery100–300 hrs
Free sample2–5 hours of audio with transcripts and full metadata, within one business day
ProvenanceCollected in Pakistan; collection date, location, contributor pseudonym, device profile and consent version on every record
LicencePerpetual, worldwide, sublicensable; derivative-model rights stated explicitly. Exclusivity optional. Licensing detail
DeliveryYour S3, GCS or Azure bucket with manifests, checksums, QA report and datasheet
Guarantee14-day acceptance window; anything outside signed tolerance re-collected free

Included in every delivery

  • Saraiki, Balochi, Brahui and Hindko speech
  • Collected in Multan, Quetta and Hazara
  • Community-network recruitment rather than crowd platforms
  • Transcription by native speakers

Built for

  • First-ever ASR coverage for Saraiki, Balochi, Brahui or Hindko
  • Extending a multilingual model to languages it currently misroutes
  • Language identification and dialect-aware routing for IVR
  • Academic and preservation work on under-documented languages

Evaluate before you buy

Start with the free sample, then a pilot

Request 2–5 hours of audio with transcripts and full metadata from Regional Languages Speech Pack — identical in format and quality to the paid corpus, no watermark, no sales call. Pilots are creditable against production and start within five business days of a signed SOW.

Average first response under 6 business hours. NDA signed same day.

  • Consent form, licence text and DPA available before you commit
  • Datasheet lists composition, QA method and known limitations
  • Subsets by language, region or condition on request
  • Custom collection to your spec using the same protocol

Answers

Questions about this dataset

Crowd platforms skew heavily toward urban, English-literate users with good bandwidth — a population that barely overlaps with native Saraiki, Balochi, Brahui or Hindko speakers. Community recruitment in Multan, Quetta and Hazara reaches the speakers a crowd platform never will.

A typical first delivery is 100–300 hours. Free sample packs, pilots and larger programmes are all available.

Yes. The pack can be licensed as a whole or as a single-language subset; ask about a subset when you request the sample.

Free samples Get a quote