New: conversational Urdu and Punjabi speech corpora now licensable  See the catalog →

Practical guides

How to run a speech data collection project

End to end: writing the spec, recruiting to quota, choosing capture conditions, transcription standards, QA gates and acceptance testing.

Key takeaways

  • Start from the deployment condition and work backwards. Every spec decision follows from what the model meets at inference.
  • Deliver a small batch in week two so the spec can be corrected while corrections are still cheap.
  • Agree the transcription convention before recording, not after. Retrofitting one is expensive.
  • Define acceptance tolerances in writing, or delivery disputes become unresolvable.

Start from the deployment condition

Before writing anything, describe precisely what the model will hear in production: which languages and dialects, on what devices, over what channel, in what acoustic environment, from which demographic, speaking in what style.

Every subsequent decision derives from that description. Teams that skip it end up specifying studio-clean audio for a product deployed to phone calls, then wondering why production accuracy does not match the benchmark.

Writing the specification

A workable spec names: target languages and dialects with minimum shares; demographic quotas with tolerances; device classes; recording environments with measured dB(A) ranges; speaking style and rate; code-switch rate if relevant; session structure and turn counts; total volume with delivery milestones; transcription convention; and acceptance tolerances.

It should also state explicitly what is not included. Most delivery disputes trace back to an assumption one party held and never wrote down.

Recruitment and consent

Recruit against quota rather than filling volume first and hoping the distribution works out. Community networks and regional partners reach people that ad networks and app-based crowdsourcing simply do not — rural speakers, older speakers, people on feature phones.

Verify identity to prevent one person operating multiple accounts, screen against quota, test for the task, and consent properly before recording anything. Disclose the pay rate in writing before the contributor accepts.

Capture conditions

Match the spec, and record the conditions as metadata rather than assuming them. Sample rate and bit depth; microphone type and distance; room characteristics; measured noise floor per session; device model.

For far-field work, capture at graded distances — 0.3 m, 1 m, 3 m, 5 m — rather than a single nominal one. For telephony, capture the actual codec path rather than downsampling wideband audio, because downsampling does not reproduce what a real GSM channel does to speech.

Transcription standards

Agree the convention before recording starts. Verbatim or cleaned? Are filled pauses transcribed? How are false starts and repairs marked? What happens with unintelligible spans? Which script for each language in code-switched speech? How are numbers, dates and acronyms written?

Then publish it to annotators and enforce it in QA. For languages with contested orthography — Saraiki, Balochi, Brahui — you must agree a house standard, because there is no external one to default to. Consistency matters more than which convention you pick.

QA gates and acceptance

Automated gates first, within minutes of upload: SNR floor, clipping, DC offset, dropout, silence ratio, duplicate fingerprint, language ID confidence, PII keyword scan.

Then human QA: a second native annotator checking transcripts blind to the first, with disagreements above threshold routed to a senior adjudicator whose rulings update a versioned rubric.

Then acceptance: a defined window against the written tolerances, with out-of-spec work re-collected at the vendor's cost and not invoiced. Without written tolerances this stage becomes an argument rather than a test.

Deliver early and often

The single highest-value process decision is delivering a small batch in week two or three rather than everything at the end.

Specs are always slightly wrong on first contact with reality. Finding that out in week two costs a spec amendment; finding out in week twelve costs a re-collection. Weekly or biweekly incremental delivery makes course correction cheap.

Related pages on this site

Answers

Frequently asked questions

A pilot of 50–200 hours typically runs 3–5 weeks including recruitment. A production programme of a few thousand hours usually runs 8–16 weeks with incremental weekly delivery throughout. New languages add 3–5 weeks of lead time to recruit, screen and calibrate a local team.

Per utterance: speaker pseudonym, dialect, region, gender, age band, device class, recording environment, measured SNR, session identifier, consent version, and the verbatim transcript with timestamps. Without this you cannot audit the distribution or weight the data during training.

It depends on the task. Scripted for TTS, wake word and constrained-grammar ASR. Spontaneous for conversational ASR, voice agents and anything where users speak freely. Spontaneous is harder and more expensive to collect and transcribe, and it is what moves real-world Word Error Rate.

Bring us a spec, or bring us a problem.

A 30-minute call with a program lead and a linguist for your language costs nothing and usually improves the spec.

Average first response: under 6 business hours. NDAs signed same day.

Free samples Get a quote