Speech & audio data · ready-made dataset
Pakistani-accented English speech dataset
More than a hundred million people in Pakistan use English as a second language, and their accents are shaped by seven different first languages. This corpus tags every speaker by L1 and proficiency so you can measure and fix exactly where an English model underperforms.
Datasheet summary
What is in Pakistani-Accented English
Read and spontaneous English graded by CEFR band, balanced by first language and region, with L1-tagged transcripts.
| Dataset ID | TDA-SP-003 |
|---|---|
| Modality | Speech |
| Languages / coverage | English, L1-tagged across 7 languages |
| Typical first delivery | 150–400 hrs |
| Free sample | 2–5 hours of audio with transcripts and full metadata, within one business day |
| Provenance | Collected in Pakistan; collection date, location, contributor pseudonym, device profile and consent version on every record |
| Licence | Perpetual, worldwide, sublicensable; derivative-model rights stated explicitly. Exclusivity optional. Licensing detail |
| Delivery | Your S3, GCS or Azure bucket with manifests, checksums, QA report and datasheet |
| Guarantee | 14-day acceptance window; anything outside signed tolerance re-collected free |
Included in every delivery
- Read and spontaneous English speech
- CEFR proficiency band per speaker
- L1 tag across seven first languages
- Balanced by first language and region
- L1-tagged transcripts
Built for
- Accent-robust English ASR for South Asian users
- Pronunciation assessment and language-learning products
- Per-L1 error analysis on an existing English model
- Speaker-adaptive TTS and voice-cloning evaluation
Evaluate before you buy
Start with the free sample, then a pilot
Request 2–5 hours of audio with transcripts and full metadata from Pakistani-Accented English — identical in format and quality to the paid corpus, no watermark, no sales call. Pilots are creditable against production and start within five business days of a signed SOW.
Average first response under 6 business hours. NDA signed same day.
- Consent form, licence text and DPA available before you commit
- Datasheet lists composition, QA method and known limitations
- Subsets by language, region or condition on request
- Custom collection to your spec using the same protocol
Answers
Questions about this dataset
Speakers are tagged by first language across seven Pakistani L1s, and the corpus is balanced by L1 and region. The datasheet lists the per-L1 speaker counts for each delivery.
Yes. Each speaker carries a CEFR band, so you can slice results by proficiency and see whether errors cluster at a particular level.
Both are included, and each recording is labelled so you can train or evaluate on either style.
Related datasets
Also in the catalog
Conversational Urdu Speech
Urdu (Karachi, Lahori, standard)
SpeechPakistani Telephony Corpus
Urdu, Punjabi, Sindhi, Pashto
SpeechRegional Languages Speech Pack
Saraiki, Balochi, Brahui, Hindko
See the full catalog · Speech & audio data services · Languages we cover