New: conversational Urdu and Punjabi speech corpora now licensable  See the catalog →

Practical guides

What AI training data actually costs

Two honest vendors can quote the same brief several times apart, and both can be right. Here is what actually moves the number, what never appears in the quote, and how to write a brief that gets you comparable bids.

Key takeaways

  • No serious vendor publishes a price list, because price is a function of the specification rather than of a catalogue.
  • Five variables move a quote more than everything else combined: spontaneity, language rarity, annotation depth, QA standard and licence scope.
  • Read speech is the cheapest audio you can buy and the least representative of what a production system hears.
  • Rework is the cost nobody budgets for. Without a re-collection clause you pay twice for data that fails acceptance.
  • Ask for a per-unit price against a written spec. A lump sum hides both the unit economics and the risk.

Why nobody publishes a price list

Buyers arrive expecting a rate card and leave frustrated that nobody has one. The frustration is fair, but the absence is not evasion. Training data is priced the way construction is priced: per unit, against a drawing. Change the drawing and the number changes, sometimes by a factor of five.

Consider a single line item, five hundred hours of Urdu speech. Read speech from prompted scripts, collected in quiet rooms, with light transcription, sits at one end. Spontaneous two-party conversation, telephone-band, verbatim transcribed, speaker-diarised, quota-balanced across age and urban or rural origin, with every contributor consented on record, sits at the other. Both are five hundred hours of Urdu. They are not the same product, and no rate card can hold both.

What a good vendor can give you before a full spec exists is a unit price with stated assumptions, and a pilot. If someone quotes a total without asking about spontaneity, quota axes or transcription standard, they have not priced your project. They have priced a project.

The five variables that move a quote

Most of the variance in any training-data quote comes from five places. Everything else is rounding.

VariableCheap endExpensive endSize of effect
SpontaneityPrompted, read from scriptUnscripted, multi-party, interrupted, emotionalLargest single factor in speech
Language and dialect rarityUrdu, Hindi, standard EnglishBalochi, Saraiki, a specific Pashto dialect groupLarge, driven by recruitment not wages
Annotation depthUtterance-level textWord-level timing, diarisation, code-switch tagging, polygonsLarge, and compounds with volume
QA standardSingle pass, spot checkDouble-blind second pass with adjudicationModerate per unit, decisive for usability
Licence scopeResearch, internal, time-limitedPerpetual, sublicensable, commercial, exclusiveModerate to large; exclusivity is the extreme

Two of these are worth dwelling on because buyers routinely misjudge them.

Rarity is a recruitment cost, not a wage cost. It is tempting to assume a language spoken in a lower-cost economy is cheaper to collect. The contributor payment is indeed smaller in absolute terms. But the expensive part of a rare-language project is finding enough qualified speakers who match your quota, and finding second reviewers competent enough to check the first ones. That search cost rises steeply as the language gets rarer, and it does not fall with local wages.

Exclusivity is the most expensive word in the contract. Asking for perpetual, sublicensable rights is normal and priced accordingly. Asking that the vendor never sell the same data to anyone else removes their ability to amortise the collection, so you absorb the whole cost of production. That is sometimes worth it. It should be a deliberate choice, not a clause your legal team pasted in.

What each data type is priced by

The unit matters as much as the rate. A price per raw hour and a price per delivered usable hour are different products, and the gap between them is your rejection rate. Insist on the delivered unit.

Data typeUsual pricing unitWhat raises the unit price
Read and prompted speechPer delivered hourRare language, speaker quotas, studio-grade capture
Spontaneous and conversational speechPer delivered hourDiarisation, verbatim transcription, multi-party audio, overlap
Telephony speechPer hour of usable audioConsent for real calls, channel separation, PII redaction
Transcription of existing audioPer audio hour or per audio minuteTimestamp granularity, speaker labels, code-switch tagging
Text corpora and parallel dataPer thousand words or per segment pairDomain expertise, human translation rather than post-edited machine output
Instruction and SFT pairsPer pairNative authorship, response length, expert or regulated domains
Preference and RLHF dataPer comparisonRater calibration, rubric complexity, raters per item
Image and video collectionPer image or per minuteRelease forms, scene diversity, device and lighting spread
Image and video annotationPer object or per framePolygons rather than boxes, class count, occlusion rules
OCR ground truthPer pageHandwriting, reading order, layout and table complexity

The costs that never appear in the quote

The line items below are real, they are yours, and they are routinely left out of the business case.

  • Rework. Data that fails your acceptance test has to be re-collected. If the contract is silent, you pay for it twice. A re-collection clause tied to a written acceptance standard is the single most valuable paragraph in a data agreement.
  • Spec drift. Halfway through collection your model team decides they also want emotion labels. That is a new product at a new price, and often a restart for the batches already recorded.
  • Your own acceptance testing. Somebody on your side has to sample, listen, measure and sign off. On a large delivery that is real engineering time, not an afternoon.
  • Legal review of provenance. If the licence chain cannot be evidenced per record, your counsel will spend weeks on it, and may still refuse the data. Cheap data with a weak consent record is the most expensive kind.
  • Redaction and handling. PII stripping, secure transfer, and any requirement to keep data inside a specific jurisdiction all carry cost, usually on both sides.
  • The gap between raw and usable. If a vendor prices raw hours and twenty per cent fails your standard, your true unit price is twenty-five per cent higher than the number on the invoice.

How to write an RFQ that gets comparable bids

Most quotes are incomparable because most briefs are underspecified, so every vendor fills the gaps differently and the cheapest bid is simply the one that assumed the least work. State the following and the bids become comparable overnight.

  • The language, and the dialect group where it matters.
  • Whether the speech is read, elicited or genuinely spontaneous.
  • The quota axes you care about and the tolerance on each: age band, gender, region, urban or rural, device.
  • The acoustic conditions you need represented, including telephone band if production is telephony.
  • The transcription standard, in writing, with a worked example of a code-switched utterance.
  • The QA standard: single pass or double-blind, and the agreement threshold you will accept.
  • The acceptance test you will run, and the consequence of failing it.
  • The licence you need: term, territory, sublicensing, exclusivity.
  • Delivery format, metadata schema and how consent evidence is delivered per record.
  • Pilot size, and the point at which pilot results either unlock or cancel the full order.

Ten lines. It takes an afternoon and it is the difference between three quotes you can compare and three quotes you cannot.

A budgeting method that survives contact with reality

Work backwards from the model, not forwards from the budget.

  • Start from the failure. Name the thing your model gets wrong in production. Accented speech, code-switching, telephone bandwidth, a script your OCR cannot segment. That failure defines the data, and most of the spec follows from it.
  • Size the smallest experiment that would move it. Usually a few tens of hours or a few thousand items, not the headline number. Buy that first.
  • Price the pilot per unit, and hold that unit price for the scale order. Get it in writing before the pilot runs, so a good pilot does not become a renegotiation.
  • Add a contingency for re-collection. Fifteen to twenty per cent of the order value is a reasonable reserve on a first engagement with a new vendor, and it should shrink on the second.
  • Measure delivered usable units. Not invoiced units. Track the ratio, and use it to compare vendors on the only number that matters.

Teams that budget this way rarely overspend. Teams that start from a number and ask what it buys almost always do, because the first thing that gets cut to fit the number is the QA standard, and the second is the consent record.

Related pages on this site

FAQ

What buyers ask about price

Because price is a function of the specification, not of a catalogue. The same 500 hours of Urdu audio can differ several times over in price depending on whether it is read or spontaneous, whether it needs verbatim transcription and diarisation, and how tightly it is quota-balanced. A published number would be wrong for almost every buyer who read it.

For a first engagement, yes. A per-unit price with a stated minimum lets you scale the order up or down without renegotiating, and it makes two vendors comparable. A fixed fee hides the unit economics, which matters when the spec changes mid-project, as it usually does.

Because the cost is recruitment and quality assurance, not the recording itself. Finding calibrated native raters for Balochi, running a second blind pass over their work and holding a demographic quota costs far more per hour than the contributor payment does. Language rarity raises the search cost, not the wage.

Large enough to expose disagreement, small enough to throw away. For speech, a few hours across every quota cell you care about. For annotation, a few hundred items double-labelled so you can measure inter-annotator agreement before scaling. A pilot that cannot fail is not a pilot.

Per unit, usually. In total, often not. Off-the-shelf data comes with someone else is distribution, someone else is consent record and no ability to fix a gap, so teams frequently buy it, discover it does not match production conditions, and commission anyway. Compare on delivered usable units, not on list price.

Send us the spec. We will quote per unit.

Assumptions listed, pilot first, unit price held for the scale order. Free samples come back within one business day.

Free samples Get a quote