New: conversational Urdu and Punjabi speech corpora now licensable  See the catalog →

Speech & audio data · ready-made dataset

Pakistani-accented English speech dataset

More than a hundred million people in Pakistan use English as a second language, and their accents are shaped by seven different first languages. This corpus tags every speaker by L1 and proficiency so you can measure and fix exactly where an English model underperforms.

Free unwatermarked samplePer-record consentTraining-rights licence14-day acceptance

Datasheet summary

What is in Pakistani-Accented English

Read and spontaneous English graded by CEFR band, balanced by first language and region, with L1-tagged transcripts.

Dataset IDTDA-SP-003
ModalitySpeech
Languages / coverageEnglish, L1-tagged across 7 languages
Typical first delivery150–400 hrs
Free sample2–5 hours of audio with transcripts and full metadata, within one business day
ProvenanceCollected in Pakistan; collection date, location, contributor pseudonym, device profile and consent version on every record
LicencePerpetual, worldwide, sublicensable; derivative-model rights stated explicitly. Exclusivity optional. Licensing detail
DeliveryYour S3, GCS or Azure bucket with manifests, checksums, QA report and datasheet
Guarantee14-day acceptance window; anything outside signed tolerance re-collected free

Included in every delivery

  • Read and spontaneous English speech
  • CEFR proficiency band per speaker
  • L1 tag across seven first languages
  • Balanced by first language and region
  • L1-tagged transcripts

Built for

  • Accent-robust English ASR for South Asian users
  • Pronunciation assessment and language-learning products
  • Per-L1 error analysis on an existing English model
  • Speaker-adaptive TTS and voice-cloning evaluation

Evaluate before you buy

Start with the free sample, then a pilot

Request 2–5 hours of audio with transcripts and full metadata from Pakistani-Accented English — identical in format and quality to the paid corpus, no watermark, no sales call. Pilots are creditable against production and start within five business days of a signed SOW.

Average first response under 6 business hours. NDA signed same day.

  • Consent form, licence text and DPA available before you commit
  • Datasheet lists composition, QA method and known limitations
  • Subsets by language, region or condition on request
  • Custom collection to your spec using the same protocol

Answers

Questions about this dataset

Speakers are tagged by first language across seven Pakistani L1s, and the corpus is balanced by L1 and region. The datasheet lists the per-L1 speaker counts for each delivery.

Yes. Each speaker carries a CEFR band, so you can slice results by proficiency and see whether errors cluster at a particular level.

Both are included, and each recording is labelled so you can train or evaluate on either style.

Free samples Get a quote