Practical guides
Free datasets vs commissioned collection: when to pay
Common Voice, FLEURS and OpenSLR are free and often sufficient. Here is exactly when open data runs out and commissioned collection becomes the cheaper option.
Key takeaways
- Use open data first. Common Voice, FLEURS and OpenSLR are free and frequently enough for a prototype.
- Open corpora share predictable limits: small volume, narrow demographics, read rather than spontaneous speech, unclear commercial licensing.
- Commission when you need a specific distribution, a defensible licence, or a domain nobody has published.
- A vendor who tells you to try the free option first is giving you a signal worth noting.
Start with the free options. Seriously.
Any data vendor who does not tell you this is optimising for their invoice rather than your outcome. Mozilla Common Voice covers a wide range of languages with permissively licensed read speech. FLEURS provides parallel speech across many languages, useful for evaluation. OpenSLR hosts a long tail of academic corpora. Hugging Face aggregates a great deal of it.
For a prototype, a feasibility study, or a first fine-tuning experiment, this is often entirely sufficient. Spending money before exhausting it is simply wasteful.
Where open data consistently runs out
Volume. Beyond the largest languages, open corpora are frequently tens of hours rather than hundreds, which is below the fine-tuning threshold for meaningful gains.
Demographics. Contributors are self-selected volunteers: skewed young, urban, male, educated, well-connected. That is a specific slice of a population, not a sample of it.
Speech style. Most open speech is read aloud from prompts. Read speech lacks the disfluency, overlap and repair of conversation, and models trained on it degrade sharply on real dialogue.
Licence clarity. Many academic corpora carry research-only or ambiguous terms. Fine for a paper; a problem when your legal team reviews a commercial deployment.
Domain. Nobody has published a corpus of your specific vertical in your specific language. That data does not exist until someone pays to create it.
The honest decision rule
Commission collection when at least one of these is true: you need a demographic or acoustic distribution that matches your actual users rather than volunteer contributors; you need a licence that will survive enterprise legal review; you need a domain or task nobody has published; or you have exhausted open data and hit a performance ceiling.
If none of those apply, keep using the free data. We will tell you that on a call, and we would rather you come back in six months with a real requirement than buy something you did not need.
The hidden cost of free data
Free is not zero-cost. Engineering time spent cleaning inconsistent formats, reconciling conflicting orthographies, filtering machine-translated contamination and auditing licences is real expenditure, and it is frequently underestimated.
For a small experiment that cost is trivially worth paying. At production scale, in a language where several partial corpora need reconciling, it can exceed the cost of commissioning clean data to a single consistent spec.
A sensible sequence
Prototype on open data. Measure where it fails and against which speaker segments. Commission a small pilot — 50 to 200 hours — targeting precisely those gaps. Re-measure. Only then scale to a production order, with a spec informed by evidence rather than assumption.
This sequence costs less overall than either extreme, and it means every unit you eventually buy is filling a gap you have actually verified exists.
Answers
Frequently asked questions
For high-resource languages it can be a strong contributor to a training mix. For low-resource languages the volumes are usually well below the fine-tuning threshold, the speakers are self-selected volunteers rather than a representative sample, and the speech is read rather than spontaneous. It is an excellent starting point and rarely a sufficient ending point.
It varies significantly. Common Voice uses CC0, which is permissive. Many academic corpora are research-only or carry ambiguous terms that were never drafted with commercial model training in mind. Check each corpus individually — do not assume a repository's default applies to everything hosted there.
It depends on language rarity, quota strictness, capture conditions, annotation depth and exclusivity — those factors move the number far more than volume does. We scope commercials per project rather than publishing a rate card, because a rate card would mislead on most of them.
Not sure whether you need commissioned data?
Tell us what you have tried and where it fell short. If open data still covers you, we will say so — it is a faster way to earn the work later.
Average first response: under 6 business hours. NDAs signed same day.