Speech & audio data · ready-made dataset
Pakistani telephony corpus
Telephony is its own acoustic domain. An 8 kHz narrowband channel discards most of the spectral information a wideband model learned from, which is why studio-trained ASR falls apart on real calls. This corpus is collected on the channel your model will actually be deployed to.
Datasheet summary
What is in Pakistani Telephony Corpus
8 kHz narrowband call-center audio, channel-separated agent/caller stems, GSM and VoIP codecs, intent and outcome labels.
| Dataset ID | TDA-SP-002 |
|---|---|
| Modality | Speech |
| Languages / coverage | Urdu, Punjabi, Sindhi, Pashto |
| Typical first delivery | 200–500 hrs |
| Free sample | 2–5 hours of audio with transcripts and full metadata, within one business day |
| Provenance | Collected in Pakistan; collection date, location, contributor pseudonym, device profile and consent version on every record |
| Licence | Perpetual, worldwide, sublicensable; derivative-model rights stated explicitly. Exclusivity optional. Licensing detail |
| Delivery | Your S3, GCS or Azure bucket with manifests, checksums, QA report and datasheet |
| Guarantee | 14-day acceptance window; anything outside signed tolerance re-collected free |
Included in every delivery
- 8 kHz narrowband call-center audio
- Channel-separated agent and caller stems
- GSM and VoIP codec conditions
- Intent and outcome labels per call
- Urdu, Punjabi, Sindhi and Pashto
Built for
- Call-center ASR and agent-assist for Pakistani-language contact centres
- IVR and voice-bot intent classification on narrowband audio
- Codec-robustness training and evaluation (GSM and VoIP)
- Conversation outcome prediction and QA scoring
Evaluate before you buy
Start with the free sample, then a pilot
Request 2–5 hours of audio with transcripts and full metadata from Pakistani Telephony Corpus — identical in format and quality to the paid corpus, no watermark, no sales call. Pilots are creditable against production and start within five business days of a signed SOW.
Average first response under 6 business hours. NDA signed same day.
- Consent form, licence text and DPA available before you commit
- Datasheet lists composition, QA method and known limitations
- Subsets by language, region or condition on request
- Custom collection to your spec using the same protocol
Answers
Questions about this dataset
Downsampling removes bandwidth but not the codec artefacts, packet loss, handset variation and background conditions of real calls. Models trained on downsampled studio speech still degrade sharply on production telephony.
Yes. Stems are channel-separated so you can train and evaluate on either side independently, or on the mixed signal.
Each call carries intent and outcome labels alongside the transcript. The label schema is documented in the datasheet that ships with every delivery and with the free sample.
Related datasets
Also in the catalog
Conversational Urdu Speech
Urdu (Karachi, Lahori, standard)
SpeechPakistani-Accented English
English, L1-tagged across 7 languages
SpeechRegional Languages Speech Pack
Saraiki, Balochi, Brahui, Hindko
See the full catalog · Speech & audio data services · Languages we cover