New: conversational Urdu and Punjabi speech corpora now licensable  See the catalog →

How we collect

The process is the product

Anyone can find people who will talk into a microphone. The difference between usable training data and expensive noise is what happens before and after — and all of it is written down, versioned and auditable.

Five stages

From a two-line brief to a signed acceptance

Each stage has an artefact you receive, and a gate we do not pass without your sign-off. No stage is a black box.

Typical end-to-end: 3–5 weeks for a pilot, 8–16 weeks for a production program, with weekly or biweekly incremental delivery throughout rather than one delivery at the end.

Specification & feasibility

A linguist for your target language joins the first call. We interrogate the spec — often arguing you out of parameters that will not help — and return a written document covering demographic quotas, dialect definitions, capture conditions, label schema, edge cases, acceptance tolerances and an explicit statement of what is not included.

Artefact: signed spec + feasibility memo

Recruitment, screening & consent

Contributors are sourced through regional partners and community networks, not ad networks. Each is ID-verified, screened against quota, tested for the task, briefed on exactly how their data will be used, and consented in their own language before producing anything. Pay rates are disclosed in writing up front.

Artefact: consent index + panel composition report

Quota-controlled collection

Field teams, studios, instrumented vehicles or our mobile capture app, depending on the spec. A live dashboard shows fill against every quota axis — language, dialect, gender, age band, region, device, environment. Out-of-quota submissions are rejected even when technically clean, because a corpus that over-samples the easy-to-reach is the exact failure we exist to prevent.

Artefact: live quota dashboard

QA, adjudication & redaction

Automated gates run within minutes of upload. Double-blind human passes run on a defined proportion. Disagreements route to a senior adjudicator whose ruling updates the versioned rubric. A PII detection and redaction pass runs before anything leaves our environment.

Artefact: per-batch QA + agreement report

Delivery & acceptance

Batches land in your bucket with manifests, checksums, per-record metadata, the QA report and the executed licence. You get a 14-day acceptance window against the signed tolerances. Anything out of tolerance is re-collected at our cost and is not invoiced.

Artefact: acceptance report + licence

Why quota sampling

500 well-distributed hours beat 5,000 convenient ones

Convenience sampling produces a corpus that mirrors who was easiest to recruit: young, urban, male, smartphone-owning, literate in the prestige dialect. Your model then performs beautifully in evaluation and fails on the users you were trying to reach.

Convenience sampling

What you get for free

  • Fast recruitment, low cost per unit
  • Heavy skew to 18–30, urban, male
  • Prestige dialect over-represented 4–6×
  • Flagship devices and good bandwidth only
  • Low acoustic and lexical variance
  • Benchmarks look great; production WER does not move
Quota-controlled sampling

What you pay a premium for

  • Distribution matches your actual user base
  • Age, gender, region and class balanced to spec
  • Dialects present in the proportions you chose
  • Device and bandwidth diversity deliberately included
  • High acoustic and lexical variance
  • Tail error — where churn lives — actually falls

The honest caveat: quota control costs more per unit and takes longer to recruit. If your target users really are urban smartphone owners in the capital city, convenience sampling is the correct and cheaper answer, and we will tell you so rather than upsell you.

Tooling

Infrastructure built for field conditions

Most annotation tooling assumes reliable broadband and a laptop. A large share of our collection happens on a mid-range Android phone with intermittent 3G.

Offline-first capture app

Records, validates and queues locally, then syncs opportunistically. Contributors in low-connectivity districts are not excluded from your corpus by their bandwidth.

Live quota dashboard

Client-facing view of fill rate against every quota axis, throughput against plan, QA pass rate and projected completion date. Updated continuously, not weekly.

Automated QA pipeline

SNR, clipping, duplication, language ID, PII scan, label-schema validation and drift detection run on every unit within minutes of upload.

Bring your own tooling

We staff on Label Studio, CVAT, Scale, Labelbox, SuperAnnotate, V7 and in-house platforms. If your schema needs a bespoke interface we build it in one to two weeks.

Secure delivery

Direct to your S3, GCS or Azure bucket with per-batch checksums and signed manifests, or via time-limited signed URLs. On-prem and VPC options for regulated work.

Weekly incremental batches

You validate the spec against real data in week two or three, not at final delivery. Course corrections are cheap early and ruinous late.

Guarantees

What we put in the contract, not just on the website

14-day acceptance

Reject anything outside signed tolerance. No invoice for rejected units.

Free re-collection

Out-of-spec work is redone at our cost, on a schedule we commit to in writing.

Named accuracy SLA

Per-task accuracy thresholds with service credits attached, measured by blind audit.

Training-rights licence

Perpetual, worldwide, sublicensable, with derivative-model rights stated explicitly.

Answers

Process questions

Kickoff call with the program lead and a linguist for your target language, a written spec draft back to you within 48 hours, and recruitment opened against draft quotas while the spec is finalised. You get dashboard access on day one, so you can see recruitment fill before any collection money is spent.

Yes, and most clients do once they see the first batch — that is why we deliberately deliver a small batch early. Changes are handled as a versioned spec amendment with a stated cost and schedule impact. Work already completed to the old spec is either accepted or re-collected by agreement, and we never silently absorb a change into quality.

ID verification at onboarding, device and location fingerprinting to catch one person operating multiple accounts, attention and comprehension checks embedded in tasks, duplicate-audio and duplicate-text fingerprinting across the whole corpus, and payment tied to QA-passed units. Fraud is a real and constant pressure in this industry; we assume it and design against it rather than being surprised by it.

Data lands in your S3, GCS or Azure bucket — or ours, with time-limited signed URLs — organised to a manifest schema we agree up front. Every batch includes SHA-256 checksums, a machine-readable manifest, per-record metadata, the QA report, the consent artefact index and the executed licence. Format conversion to your loader's expectations is included, not extra.

Bring us a spec, or bring us a problem.

Either works. A 30-minute call with a program lead and a linguist for your language costs nothing and usually changes the spec.

Average first response: under 6 business hours. NDAs signed same day.

Free samples Get a quote