Speech & audio data
Speech data from the people your model actually serves
Scripted, spontaneous, conversational, far-field and telephony audio across every major language of Pakistan and dialects — recorded to a demographic and acoustic spec you approve, transcribed by natives, and licensed for training.
Collection types
Every kind of speech, collected the way the task demands
The right collection design depends entirely on what the model has to do at inference. We start from the deployment condition and work backwards.
Scripted & prompted speech
Phonetically balanced prompts, digit strings, addresses, names, commands and domain phrases. Best for TTS, wake word and constrained-grammar ASR.
Spontaneous & conversational
Unscripted monologue, two-party dialogue, interviews and role-play with natural disfluency, overlap, laughter and repair. The data that fixes real-world WER.
Telephony & call center
8 kHz narrowband, GSM and VoIP codecs, hold music, IVR barge-in and channel-separated two-party calls with agent and caller stems.
Far-field & in-device
Multi-mic arrays at graded distances, in-car at speed, smart-speaker and kitchen conditions, with measured room impulse and noise floor.
Speaker & voice biometrics
Multi-session enrolment for speaker verification, with consented re-recording windows and controlled channel variation across sessions.
Accented & L2 English
English as spoken by Urdu, Punjabi, Sindhi and Pashto natives, graded by CEFR band and tagged by first language — heavily requested and barely available anywhere.
Specification control
You control the distribution, not just the volume
Volume without distribution control produces confident models that fail on the tail. Every axis below is a quota we recruit against and reject out-of-spec submissions on.
- Language, dialect and regional variety, down to district level
- Gender balance, age bands (13–17 with guardian consent, 18–29, 30–49, 50–64, 65+)
- Urban / peri-urban / rural split and socio-economic band
- Device class: flagship, mid-range Android, feature phone, headset, array mic
- Recording environment and measured noise level in dB(A)
- Speaking style, speech rate, emotional register and code-switch rate
- Session count per speaker and inter-session interval for biometrics
Example: conversational Urdu, 1,000 hrs
Typical delivery specification — all fields are configurable.
Applications
What teams build with it
| Use case | What we collect | Typical scale |
|---|---|---|
| Multilingual ASR | Spontaneous + telephony audio, verbatim transcripts, forced alignment | 1,000–20,000 hrs per language |
| Expressive TTS | Studio-grade single-speaker, 4 emotional registers, phonetic balance | 20–60 hrs per voice |
| Voice agents / IVR | Two-party role-play, barge-in, noise, code-switching | 300–2,000 hrs |
| Wake word | Positive triggers at graded distance + hard negatives | 50k–500k utterances |
| Speaker verification | Multi-session enrolment across channels and 6-month intervals | 5k–50k speakers |
| Speech translation | Aligned source audio + native target transcripts | 500–5,000 hrs per pair |
| Emotion & intent | Elicited and natural affect, labeled by 3 native raters | 100–800 hrs |
Quality
How we know the audio is good before you do
Automated gates run on every file within minutes of upload. Human QA runs on a stratified sample of every batch. Nothing reaches you unaudited.
- Automated: SNR floor, clipping, DC offset, dropout, silence ratio, duplicate fingerprint
- Automated: language ID confidence, speaker-change detection, PII keyword scan
- Human: verbatim transcript check by a second native annotator, blind to the first
- Human: adjudication by a senior linguist on any disagreement above threshold
- Reported: inter-annotator agreement, WER of transcripts against gold subset, quota fill by axis
- Delivered: per-batch QA report you can reject on, with 14-day acceptance window
Answers
Speech data questions
Default is 48 kHz / 16-bit mono WAV with a separate lossless archive, but we match your pipeline: 16 kHz for telephony models, 8 kHz µ-law for IVR parity, Opus or FLAC for storage-constrained delivery. Each utterance ships with a JSON sidecar containing speaker pseudonym, demographics, device profile, environment tag, SNR measurement and verbatim transcript.
Yes, with two routes. Where the client owns the call recordings we operate as processor and handle transcription, diarization and redaction inside your environment. Where you need net-new audio we run scripted two-party and three-party role-play collection with trained agents, which produces the turn-taking, overlap and hold-music artefacts that scripted single-speaker data never captures.
We treat it as a first-class spec parameter, not noise. You set a target code-switch rate (for example, 30–40% of Urdu utterances containing at least one English content word), and recruitment plus prompt design are built to hit it. Transcripts use a documented tagging convention so the switch points are recoverable for training.
Yes. We field multi-microphone arrays at set distances (0.3 m / 1 m / 3 m / 5 m), in-vehicle capture at defined speeds and window states, and controlled-noise environments — market, street, kitchen, factory floor, mosque courtyard, monsoon rain — with measured dB(A) levels recorded per session.
Standard SLA is ≥98% word accuracy on verbatim transcription for high-resource languages and ≥95% for low-resource languages where orthography is contested, measured by third-sample blind audit. Where a language has no settled written standard we agree a house orthography with you up front and enforce it — consistency matters more than any single convention.
Send us the acoustic condition your model keeps failing on.
We will tell you what collection design fixes it, how many hours it realistically takes, and what it costs. Free 20-minute technical scoping with a program lead, not a salesperson.
Average first response: under 6 business hours. NDAs signed same day.