New: conversational Urdu and Punjabi speech corpora now licensable  See the catalog →

Speech & audio data · ready-made dataset

Conversational Urdu speech dataset

Spontaneous two-party Urdu conversation — the kind of audio that voice agents, call analytics and meeting tools actually encounter in production, and the kind that read-speech corpora never contain.

Free unwatermarked samplePer-record consentTraining-rights licence14-day acceptance

Datasheet summary

What is in Conversational Urdu Speech

Two-party spontaneous telephone and in-person conversation, verbatim transcripts, speaker diarization, balanced across 4 age bands and urban/rural.

Dataset IDTDA-SP-001
ModalitySpeech
Languages / coverageUrdu (Karachi, Lahori, standard)
Typical first delivery200–500 hrs
Free sample2–5 hours of audio with transcripts and full metadata, within one business day
ProvenanceCollected in Pakistan; collection date, location, contributor pseudonym, device profile and consent version on every record
LicencePerpetual, worldwide, sublicensable; derivative-model rights stated explicitly. Exclusivity optional. Licensing detail
DeliveryYour S3, GCS or Azure bucket with manifests, checksums, QA report and datasheet
Guarantee14-day acceptance window; anything outside signed tolerance re-collected free

Included in every delivery

  • Two-party spontaneous dialogue, telephone and in-person
  • Verbatim transcripts with speaker diarization
  • Balanced across four age bands and urban/rural
  • Karachi, Lahori and standard Urdu varieties represented

Built for

  • Fine-tuning Whisper, MMS or wav2vec 2.0 checkpoints for conversational Urdu ASR
  • Diarization and turn-taking models for Urdu voice agents and IVR
  • Call analytics, summarisation and intent detection on real dialogue
  • Building a held-out test set that reflects real users rather than studio speakers

Evaluate before you buy

Start with the free sample, then a pilot

Request 2–5 hours of audio with transcripts and full metadata from Conversational Urdu Speech — identical in format and quality to the paid corpus, no watermark, no sales call. Pilots are creditable against production and start within five business days of a signed SOW.

Average first response under 6 business hours. NDA signed same day.

  • Consent form, licence text and DPA available before you commit
  • Datasheet lists composition, QA method and known limitations
  • Subsets by language, region or condition on request
  • Custom collection to your spec using the same protocol

Answers

Questions about this dataset

Spontaneous. Contributors hold real two-party conversations; nothing is scripted or read from prompts. Transcripts are verbatim, so disfluencies, repairs and overlap are preserved rather than cleaned away.

Recruitment is quota-controlled across four age bands and an urban/rural split, with Karachi, Lahori and standard Urdu varieties represented. The datasheet lists the exact composition of each delivery.

Yes. Request a free, unwatermarked sample pack — typically 2–5 hours of audio with transcripts and the full datasheet — delivered within one business day, no sales call required.

Free samples Get a quote