How we collect
The process is the product
Anyone can find people who will talk into a microphone. The difference between usable training data and expensive noise is what happens before and after — and all of it is written down, versioned and auditable.
Five stages
From a two-line brief to a signed acceptance
Each stage has an artefact you receive, and a gate we do not pass without your sign-off. No stage is a black box.
Typical end-to-end: 3–5 weeks for a pilot, 8–16 weeks for a production program, with weekly or biweekly incremental delivery throughout rather than one delivery at the end.
Specification & feasibility
A linguist for your target language joins the first call. We interrogate the spec — often arguing you out of parameters that will not help — and return a written document covering demographic quotas, dialect definitions, capture conditions, label schema, edge cases, acceptance tolerances and an explicit statement of what is not included.
Recruitment, screening & consent
Contributors are sourced through regional partners and community networks, not ad networks. Each is ID-verified, screened against quota, tested for the task, briefed on exactly how their data will be used, and consented in their own language before producing anything. Pay rates are disclosed in writing up front.
Quota-controlled collection
Field teams, studios, instrumented vehicles or our mobile capture app, depending on the spec. A live dashboard shows fill against every quota axis — language, dialect, gender, age band, region, device, environment. Out-of-quota submissions are rejected even when technically clean, because a corpus that over-samples the easy-to-reach is the exact failure we exist to prevent.
QA, adjudication & redaction
Automated gates run within minutes of upload. Double-blind human passes run on a defined proportion. Disagreements route to a senior adjudicator whose ruling updates the versioned rubric. A PII detection and redaction pass runs before anything leaves our environment.
Delivery & acceptance
Batches land in your bucket with manifests, checksums, per-record metadata, the QA report and the executed licence. You get a 14-day acceptance window against the signed tolerances. Anything out of tolerance is re-collected at our cost and is not invoiced.
Why quota sampling
500 well-distributed hours beat 5,000 convenient ones
Convenience sampling produces a corpus that mirrors who was easiest to recruit: young, urban, male, smartphone-owning, literate in the prestige dialect. Your model then performs beautifully in evaluation and fails on the users you were trying to reach.
What you get for free
- Fast recruitment, low cost per unit
- Heavy skew to 18–30, urban, male
- Prestige dialect over-represented 4–6×
- Flagship devices and good bandwidth only
- Low acoustic and lexical variance
- Benchmarks look great; production WER does not move
What you pay a premium for
- Distribution matches your actual user base
- Age, gender, region and class balanced to spec
- Dialects present in the proportions you chose
- Device and bandwidth diversity deliberately included
- High acoustic and lexical variance
- Tail error — where churn lives — actually falls
The honest caveat: quota control costs more per unit and takes longer to recruit. If your target users really are urban smartphone owners in the capital city, convenience sampling is the correct and cheaper answer, and we will tell you so rather than upsell you.
Tooling
Infrastructure built for field conditions
Most annotation tooling assumes reliable broadband and a laptop. A large share of our collection happens on a mid-range Android phone with intermittent 3G.
Offline-first capture app
Records, validates and queues locally, then syncs opportunistically. Contributors in low-connectivity districts are not excluded from your corpus by their bandwidth.
Live quota dashboard
Client-facing view of fill rate against every quota axis, throughput against plan, QA pass rate and projected completion date. Updated continuously, not weekly.
Automated QA pipeline
SNR, clipping, duplication, language ID, PII scan, label-schema validation and drift detection run on every unit within minutes of upload.
Bring your own tooling
We staff on Label Studio, CVAT, Scale, Labelbox, SuperAnnotate, V7 and in-house platforms. If your schema needs a bespoke interface we build it in one to two weeks.
Secure delivery
Direct to your S3, GCS or Azure bucket with per-batch checksums and signed manifests, or via time-limited signed URLs. On-prem and VPC options for regulated work.
Weekly incremental batches
You validate the spec against real data in week two or three, not at final delivery. Course corrections are cheap early and ruinous late.
Guarantees
What we put in the contract, not just on the website
14-day acceptance
Reject anything outside signed tolerance. No invoice for rejected units.
Free re-collection
Out-of-spec work is redone at our cost, on a schedule we commit to in writing.
Named accuracy SLA
Per-task accuracy thresholds with service credits attached, measured by blind audit.
Training-rights licence
Perpetual, worldwide, sublicensable, with derivative-model rights stated explicitly.
Answers
Process questions
Kickoff call with the program lead and a linguist for your target language, a written spec draft back to you within 48 hours, and recruitment opened against draft quotas while the spec is finalised. You get dashboard access on day one, so you can see recruitment fill before any collection money is spent.
Yes, and most clients do once they see the first batch — that is why we deliberately deliver a small batch early. Changes are handled as a versioned spec amendment with a stated cost and schedule impact. Work already completed to the old spec is either accepted or re-collected by agreement, and we never silently absorb a change into quality.
ID verification at onboarding, device and location fingerprinting to catch one person operating multiple accounts, attention and comprehension checks embedded in tasks, duplicate-audio and duplicate-text fingerprinting across the whole corpus, and payment tied to QA-passed units. Fraud is a real and constant pressure in this industry; we assume it and design against it rather than being surprised by it.
Data lands in your S3, GCS or Azure bucket — or ours, with time-limited signed URLs — organised to a manifest schema we agree up front. Every batch includes SHA-256 checksums, a machine-readable manifest, per-record metadata, the QA report, the consent artefact index and the executed licence. Format conversion to your loader's expectations is included, not extra.
Bring us a spec, or bring us a problem.
Either works. A 30-minute call with a program lead and a linguist for your language costs nothing and usually changes the spec.
Average first response: under 6 business hours. NDAs signed same day.