New: conversational Urdu and Punjabi speech corpora now licensable  See the catalog →

Practical guides

Quota sampling beats volume: how to specify a dataset

Why 500 well-distributed hours beat 5,000 convenience-sampled ones, which axes to quota on, and how to write a data specification that survives contact with reality.

Key takeaways

  • Convenience sampling collects whoever is easiest to reach, which is systematically the wrong people.
  • Quota sampling defines the target distribution first and rejects out-of-quota submissions even when they are clean.
  • The axes that matter: dialect, region, gender, age band, urban/rural, device class, environment, speaking style.
  • Quota control costs more per unit and reduces tail error, which is where production failures actually live.

The default failure mode

Left unspecified, data collection collects whoever volunteers. In practice that is consistently young, urban, male, more educated, better connected and using a better phone than the population you are actually serving.

The resulting model performs well for that group and measurably worse for everyone else. Because the evaluation set is usually drawn from the same pool, nothing in your metrics reveals the problem until users do.

What quota sampling does differently

You define the target distribution before recruiting, then recruit against it and enforce it. Concretely: 50/50 gender split within a stated tolerance, minimum percentages per age band, specified urban and rural proportions, named dialects with minimum shares, a device mix spanning flagship to budget handsets, and environment conditions with measured noise levels.

The enforcement part matters most. Submissions that fall outside quota get rejected even when they are technically clean, because accepting them is exactly how a corpus drifts back toward convenience sampling one reasonable exception at a time.

The axes worth quotaing on

For speech: language and dialect down to district level where it matters; gender; age bands; urban, peri-urban and rural split; socio-economic band; device class; recording environment with measured dB(A); speaking style and rate; code-switch rate; and session count per speaker where you need speaker-verification data.

You will not quota on all of them. Pick the four or five that plausibly affect your model's behaviour for your users, and specify those tightly rather than specifying everything loosely.

Writing a specification that survives contact

A usable spec states tolerances, not just targets — “50/50 gender, ±3%” is actionable where “balanced” is not. It names edge cases explicitly. It defines what constitutes a rejected unit. And it states what is not included, which prevents the most common category of delivery dispute.

It should also survive being read by someone who was not in the kick-off call, because in a twelve-week programme that person will exist.

The honest trade-off

Quota control costs more per unit and takes longer to recruit. Finding a 55-year-old rural Saraiki speaker with a budget handset is genuinely harder than finding another Lahore university student.

It is worth paying when your users include that person. If your product genuinely only serves urban smartphone owners in major cities, convenience sampling is the correct and cheaper answer, and a good vendor will tell you so rather than upselling rigour you do not need.

Related pages on this site

Answers

Frequently asked questions

Defining a target demographic and technical distribution before recruiting, then recruiting against those quotas and rejecting submissions that fall outside them. It contrasts with convenience sampling, which accepts whoever is easiest to reach and systematically over-represents them.

No, and this is the most consistently expensive misconception in dataset procurement. Additional data drawn from a distribution you already over-represent produces diminishing returns. Data filling a genuine gap produces disproportionate ones. Well-distributed hundreds routinely outperform badly-distributed thousands.

Work backwards from deployment. List the ways your actual users differ from each other in ways that plausibly change the input signal — accent, device, environment, age, register. Those are your axes. If you cannot describe your users on an axis, you probably should not be quotaing on it.

Bring us a spec, or bring us a problem.

Either works. A 30-minute call with a program lead and a linguist costs nothing and usually improves the spec.

Average first response: under 6 business hours. NDAs signed same day.

Free samples Get a quote