Practical guides
What AI training data actually costs
Two honest vendors can quote the same brief several times apart, and both can be right. Here is what actually moves the number, what never appears in the quote, and how to write a brief that gets you comparable bids.
Key takeaways
- No serious vendor publishes a price list, because price is a function of the specification rather than of a catalogue.
- Five variables move a quote more than everything else combined: spontaneity, language rarity, annotation depth, QA standard and licence scope.
- Read speech is the cheapest audio you can buy and the least representative of what a production system hears.
- Rework is the cost nobody budgets for. Without a re-collection clause you pay twice for data that fails acceptance.
- Ask for a per-unit price against a written spec. A lump sum hides both the unit economics and the risk.
Why nobody publishes a price list
Buyers arrive expecting a rate card and leave frustrated that nobody has one. The frustration is fair, but the absence is not evasion. Training data is priced the way construction is priced: per unit, against a drawing. Change the drawing and the number changes, sometimes by a factor of five.
Consider a single line item, five hundred hours of Urdu speech. Read speech from prompted scripts, collected in quiet rooms, with light transcription, sits at one end. Spontaneous two-party conversation, telephone-band, verbatim transcribed, speaker-diarised, quota-balanced across age and urban or rural origin, with every contributor consented on record, sits at the other. Both are five hundred hours of Urdu. They are not the same product, and no rate card can hold both.
What a good vendor can give you before a full spec exists is a unit price with stated assumptions, and a pilot. If someone quotes a total without asking about spontaneity, quota axes or transcription standard, they have not priced your project. They have priced a project.
The five variables that move a quote
Most of the variance in any training-data quote comes from five places. Everything else is rounding.
| Variable | Cheap end | Expensive end | Size of effect |
|---|---|---|---|
| Spontaneity | Prompted, read from script | Unscripted, multi-party, interrupted, emotional | Largest single factor in speech |
| Language and dialect rarity | Urdu, Hindi, standard English | Balochi, Saraiki, a specific Pashto dialect group | Large, driven by recruitment not wages |
| Annotation depth | Utterance-level text | Word-level timing, diarisation, code-switch tagging, polygons | Large, and compounds with volume |
| QA standard | Single pass, spot check | Double-blind second pass with adjudication | Moderate per unit, decisive for usability |
| Licence scope | Research, internal, time-limited | Perpetual, sublicensable, commercial, exclusive | Moderate to large; exclusivity is the extreme |
Two of these are worth dwelling on because buyers routinely misjudge them.
Rarity is a recruitment cost, not a wage cost. It is tempting to assume a language spoken in a lower-cost economy is cheaper to collect. The contributor payment is indeed smaller in absolute terms. But the expensive part of a rare-language project is finding enough qualified speakers who match your quota, and finding second reviewers competent enough to check the first ones. That search cost rises steeply as the language gets rarer, and it does not fall with local wages.
Exclusivity is the most expensive word in the contract. Asking for perpetual, sublicensable rights is normal and priced accordingly. Asking that the vendor never sell the same data to anyone else removes their ability to amortise the collection, so you absorb the whole cost of production. That is sometimes worth it. It should be a deliberate choice, not a clause your legal team pasted in.
What each data type is priced by
The unit matters as much as the rate. A price per raw hour and a price per delivered usable hour are different products, and the gap between them is your rejection rate. Insist on the delivered unit.
| Data type | Usual pricing unit | What raises the unit price |
|---|---|---|
| Read and prompted speech | Per delivered hour | Rare language, speaker quotas, studio-grade capture |
| Spontaneous and conversational speech | Per delivered hour | Diarisation, verbatim transcription, multi-party audio, overlap |
| Telephony speech | Per hour of usable audio | Consent for real calls, channel separation, PII redaction |
| Transcription of existing audio | Per audio hour or per audio minute | Timestamp granularity, speaker labels, code-switch tagging |
| Text corpora and parallel data | Per thousand words or per segment pair | Domain expertise, human translation rather than post-edited machine output |
| Instruction and SFT pairs | Per pair | Native authorship, response length, expert or regulated domains |
| Preference and RLHF data | Per comparison | Rater calibration, rubric complexity, raters per item |
| Image and video collection | Per image or per minute | Release forms, scene diversity, device and lighting spread |
| Image and video annotation | Per object or per frame | Polygons rather than boxes, class count, occlusion rules |
| OCR ground truth | Per page | Handwriting, reading order, layout and table complexity |
The costs that never appear in the quote
The line items below are real, they are yours, and they are routinely left out of the business case.
- Rework. Data that fails your acceptance test has to be re-collected. If the contract is silent, you pay for it twice. A re-collection clause tied to a written acceptance standard is the single most valuable paragraph in a data agreement.
- Spec drift. Halfway through collection your model team decides they also want emotion labels. That is a new product at a new price, and often a restart for the batches already recorded.
- Your own acceptance testing. Somebody on your side has to sample, listen, measure and sign off. On a large delivery that is real engineering time, not an afternoon.
- Legal review of provenance. If the licence chain cannot be evidenced per record, your counsel will spend weeks on it, and may still refuse the data. Cheap data with a weak consent record is the most expensive kind.
- Redaction and handling. PII stripping, secure transfer, and any requirement to keep data inside a specific jurisdiction all carry cost, usually on both sides.
- The gap between raw and usable. If a vendor prices raw hours and twenty per cent fails your standard, your true unit price is twenty-five per cent higher than the number on the invoice.
How to write an RFQ that gets comparable bids
Most quotes are incomparable because most briefs are underspecified, so every vendor fills the gaps differently and the cheapest bid is simply the one that assumed the least work. State the following and the bids become comparable overnight.
- The language, and the dialect group where it matters.
- Whether the speech is read, elicited or genuinely spontaneous.
- The quota axes you care about and the tolerance on each: age band, gender, region, urban or rural, device.
- The acoustic conditions you need represented, including telephone band if production is telephony.
- The transcription standard, in writing, with a worked example of a code-switched utterance.
- The QA standard: single pass or double-blind, and the agreement threshold you will accept.
- The acceptance test you will run, and the consequence of failing it.
- The licence you need: term, territory, sublicensing, exclusivity.
- Delivery format, metadata schema and how consent evidence is delivered per record.
- Pilot size, and the point at which pilot results either unlock or cancel the full order.
Ten lines. It takes an afternoon and it is the difference between three quotes you can compare and three quotes you cannot.
A budgeting method that survives contact with reality
Work backwards from the model, not forwards from the budget.
- Start from the failure. Name the thing your model gets wrong in production. Accented speech, code-switching, telephone bandwidth, a script your OCR cannot segment. That failure defines the data, and most of the spec follows from it.
- Size the smallest experiment that would move it. Usually a few tens of hours or a few thousand items, not the headline number. Buy that first.
- Price the pilot per unit, and hold that unit price for the scale order. Get it in writing before the pilot runs, so a good pilot does not become a renegotiation.
- Add a contingency for re-collection. Fifteen to twenty per cent of the order value is a reasonable reserve on a first engagement with a new vendor, and it should shrink on the second.
- Measure delivered usable units. Not invoiced units. Track the ratio, and use it to compare vendors on the only number that matters.
Teams that budget this way rarely overspend. Teams that start from a number and ask what it buys almost always do, because the first thing that gets cut to fit the number is the QA standard, and the second is the consent record.
FAQ
What buyers ask about price
Because price is a function of the specification, not of a catalogue. The same 500 hours of Urdu audio can differ several times over in price depending on whether it is read or spontaneous, whether it needs verbatim transcription and diarisation, and how tightly it is quota-balanced. A published number would be wrong for almost every buyer who read it.
For a first engagement, yes. A per-unit price with a stated minimum lets you scale the order up or down without renegotiating, and it makes two vendors comparable. A fixed fee hides the unit economics, which matters when the spec changes mid-project, as it usually does.
Because the cost is recruitment and quality assurance, not the recording itself. Finding calibrated native raters for Balochi, running a second blind pass over their work and holding a demographic quota costs far more per hour than the contributor payment does. Language rarity raises the search cost, not the wage.
Large enough to expose disagreement, small enough to throw away. For speech, a few hours across every quota cell you care about. For annotation, a few hundred items double-labelled so you can measure inter-annotator agreement before scaling. A pilot that cannot fail is not a pilot.
Per unit, usually. In total, often not. Off-the-shelf data comes with someone else is distribution, someone else is consent record and no ability to fix a gap, so teams frequently buy it, discover it does not match production conditions, and commission anyway. Compare on delivered usable units, not on list price.
Send us the spec. We will quote per unit.
Assumptions listed, pilot first, unit price held for the scale order. Free samples come back within one business day.