New: conversational Urdu and Punjabi speech corpora now licensable  See the catalog →

Speech & audio data

Speech data from the people your model actually serves

Scripted, spontaneous, conversational, far-field and telephony audio across every major language of Pakistan and dialects — recorded to a demographic and acoustic spec you approve, transcribed by natives, and licensed for training.

ASR & TTSWake wordSpeaker IDDiarizationEmotionAccented English
0
Languages & major dialects
0%
Verbatim transcript accuracy, contracted
0 days
Acceptance window on every batch
0
Cost to you for out-of-spec re-collection

Collection types

Every kind of speech, collected the way the task demands

The right collection design depends entirely on what the model has to do at inference. We start from the deployment condition and work backwards.

Scripted & prompted speech

Phonetically balanced prompts, digit strings, addresses, names, commands and domain phrases. Best for TTS, wake word and constrained-grammar ASR.

Spontaneous & conversational

Unscripted monologue, two-party dialogue, interviews and role-play with natural disfluency, overlap, laughter and repair. The data that fixes real-world WER.

Telephony & call center

8 kHz narrowband, GSM and VoIP codecs, hold music, IVR barge-in and channel-separated two-party calls with agent and caller stems.

Far-field & in-device

Multi-mic arrays at graded distances, in-car at speed, smart-speaker and kitchen conditions, with measured room impulse and noise floor.

Speaker & voice biometrics

Multi-session enrolment for speaker verification, with consented re-recording windows and controlled channel variation across sessions.

Accented & L2 English

English as spoken by Urdu, Punjabi, Sindhi and Pashto natives, graded by CEFR band and tagged by first language — heavily requested and barely available anywhere.

Specification control

You control the distribution, not just the volume

Volume without distribution control produces confident models that fail on the tail. Every axis below is a quota we recruit against and reject out-of-spec submissions on.

  • Language, dialect and regional variety, down to district level
  • Gender balance, age bands (13–17 with guardian consent, 18–29, 30–49, 50–64, 65+)
  • Urban / peri-urban / rural split and socio-economic band
  • Device class: flagship, mid-range Android, feature phone, headset, array mic
  • Recording environment and measured noise level in dB(A)
  • Speaking style, speech rate, emotional register and code-switch rate
  • Session count per speaker and inter-session interval for biometrics

Example: conversational Urdu, 1,000 hrs

Typical delivery specification — all fields are configurable.

Sample rate48 kHz / 16-bit
Speakers1,200 unique
Gender split50 / 50 ±3%
Age bands4 bands, min 18% each
ProvincesPunjab, Sindh, KP, Balochistan
DialectsStandard, Karachi, Lahori
Environment60% quiet, 40% ambient
Noise floor< 35 dB(A) / 45–65 dB(A)
Turns per session40–90
TranscriptionVerbatim + timestamps
DeliveryWeekly batches, 12 weeks

Applications

What teams build with it

Use caseWhat we collectTypical scale
Multilingual ASRSpontaneous + telephony audio, verbatim transcripts, forced alignment1,000–20,000 hrs per language
Expressive TTSStudio-grade single-speaker, 4 emotional registers, phonetic balance20–60 hrs per voice
Voice agents / IVRTwo-party role-play, barge-in, noise, code-switching300–2,000 hrs
Wake wordPositive triggers at graded distance + hard negatives50k–500k utterances
Speaker verificationMulti-session enrolment across channels and 6-month intervals5k–50k speakers
Speech translationAligned source audio + native target transcripts500–5,000 hrs per pair
Emotion & intentElicited and natural affect, labeled by 3 native raters100–800 hrs

Quality

How we know the audio is good before you do

Automated gates run on every file within minutes of upload. Human QA runs on a stratified sample of every batch. Nothing reaches you unaudited.

  • Automated: SNR floor, clipping, DC offset, dropout, silence ratio, duplicate fingerprint
  • Automated: language ID confidence, speaker-change detection, PII keyword scan
  • Human: verbatim transcript check by a second native annotator, blind to the first
  • Human: adjudication by a senior linguist on any disagreement above threshold
  • Reported: inter-annotator agreement, WER of transcripts against gold subset, quota fill by axis
  • Delivered: per-batch QA report you can reject on, with 14-day acceptance window

Answers

Speech data questions

Default is 48 kHz / 16-bit mono WAV with a separate lossless archive, but we match your pipeline: 16 kHz for telephony models, 8 kHz µ-law for IVR parity, Opus or FLAC for storage-constrained delivery. Each utterance ships with a JSON sidecar containing speaker pseudonym, demographics, device profile, environment tag, SNR measurement and verbatim transcript.

Yes, with two routes. Where the client owns the call recordings we operate as processor and handle transcription, diarization and redaction inside your environment. Where you need net-new audio we run scripted two-party and three-party role-play collection with trained agents, which produces the turn-taking, overlap and hold-music artefacts that scripted single-speaker data never captures.

We treat it as a first-class spec parameter, not noise. You set a target code-switch rate (for example, 30–40% of Urdu utterances containing at least one English content word), and recruitment plus prompt design are built to hit it. Transcripts use a documented tagging convention so the switch points are recoverable for training.

Yes. We field multi-microphone arrays at set distances (0.3 m / 1 m / 3 m / 5 m), in-vehicle capture at defined speeds and window states, and controlled-noise environments — market, street, kitchen, factory floor, mosque courtyard, monsoon rain — with measured dB(A) levels recorded per session.

Standard SLA is ≥98% word accuracy on verbatim transcription for high-resource languages and ≥95% for low-resource languages where orthography is contested, measured by third-sample blind audit. Where a language has no settled written standard we agree a house orthography with you up front and enforce it — consistency matters more than any single convention.

Send us the acoustic condition your model keeps failing on.

We will tell you what collection design fixes it, how many hours it realistically takes, and what it costs. Free 20-minute technical scoping with a program lead, not a salesperson.

Average first response: under 6 business hours. NDAs signed same day.

Free samples Get a quote