Speech & audio data · ready-made dataset
Conversational Urdu speech dataset
Spontaneous two-party Urdu conversation — the kind of audio that voice agents, call analytics and meeting tools actually encounter in production, and the kind that read-speech corpora never contain.
Datasheet summary
What is in Conversational Urdu Speech
Two-party spontaneous telephone and in-person conversation, verbatim transcripts, speaker diarization, balanced across 4 age bands and urban/rural.
| Dataset ID | TDA-SP-001 |
|---|---|
| Modality | Speech |
| Languages / coverage | Urdu (Karachi, Lahori, standard) |
| Typical first delivery | 200–500 hrs |
| Free sample | 2–5 hours of audio with transcripts and full metadata, within one business day |
| Provenance | Collected in Pakistan; collection date, location, contributor pseudonym, device profile and consent version on every record |
| Licence | Perpetual, worldwide, sublicensable; derivative-model rights stated explicitly. Exclusivity optional. Licensing detail |
| Delivery | Your S3, GCS or Azure bucket with manifests, checksums, QA report and datasheet |
| Guarantee | 14-day acceptance window; anything outside signed tolerance re-collected free |
Included in every delivery
- Two-party spontaneous dialogue, telephone and in-person
- Verbatim transcripts with speaker diarization
- Balanced across four age bands and urban/rural
- Karachi, Lahori and standard Urdu varieties represented
Built for
- Fine-tuning Whisper, MMS or wav2vec 2.0 checkpoints for conversational Urdu ASR
- Diarization and turn-taking models for Urdu voice agents and IVR
- Call analytics, summarisation and intent detection on real dialogue
- Building a held-out test set that reflects real users rather than studio speakers
Evaluate before you buy
Start with the free sample, then a pilot
Request 2–5 hours of audio with transcripts and full metadata from Conversational Urdu Speech — identical in format and quality to the paid corpus, no watermark, no sales call. Pilots are creditable against production and start within five business days of a signed SOW.
Average first response under 6 business hours. NDA signed same day.
- Consent form, licence text and DPA available before you commit
- Datasheet lists composition, QA method and known limitations
- Subsets by language, region or condition on request
- Custom collection to your spec using the same protocol
Answers
Questions about this dataset
Spontaneous. Contributors hold real two-party conversations; nothing is scripted or read from prompts. Transcripts are verbatim, so disfluencies, repairs and overlap are preserved rather than cleaned away.
Recruitment is quota-controlled across four age bands and an urban/rural split, with Karachi, Lahori and standard Urdu varieties represented. The datasheet lists the exact composition of each delivery.
Yes. Request a free, unwatermarked sample pack — typically 2–5 hours of audio with transcripts and the full datasheet — delivered within one business day, no sales call required.
Related datasets
Also in the catalog
Pakistani Telephony Corpus
Urdu, Punjabi, Sindhi, Pashto
SpeechPakistani-Accented English
English, L1-tagged across 7 languages
SpeechRegional Languages Speech Pack
Saraiki, Balochi, Brahui, Hindko
See the full catalog · Speech & audio data services · Languages we cover