Legal & ethics
AI training data licensing: what your contract must actually say
What a training-rights licence must contain, why most content licences do not cover model training, and the clauses that stop deals in enterprise legal review.
Key takeaways
- Most content licences predate machine learning and do not clearly grant training rights. Silence is not permission.
- A usable licence must name training, fine-tuning, evaluation and derivative model distribution explicitly.
- Perpetual and irrevocable matter more than they sound: a licence that can be withdrawn is a model you may have to retrain.
- Provenance documentation is what turns a licence from a promise into something auditable.
Why most data licences do not cover model training
Nearly every standard content licence was drafted for reproduction, display and distribution — the uses that existed when the template was written. Training a model is none of those. It is a statistical extraction that produces an artefact bearing no direct copy of the input.
That novelty cuts both ways. It means older licences rarely prohibit training explicitly, and it means they rarely permit it explicitly either. Enterprise counsel reviewing a nine-figure model programme will not accept ambiguity, so the absence of a prohibition is not the same as a grant.
The five clauses that actually matter
Scope of use. The licence must name machine-learning training, fine-tuning and evaluation as permitted uses in those words. Generic “any lawful purpose” language is weaker than it looks under scrutiny.
Derivative model rights. A model trained on the data is a derivative artefact. If the licence does not address whether you may distribute it, you have a gap precisely where your commercial value sits.
Perpetual and irrevocable. A licence that expires or can be withdrawn creates an obligation you cannot practically discharge — you cannot un-train a model. Perpetuity is not a nice-to-have.
Sublicensability. If you ship a model to customers, or embed it in a product, you are effectively passing rights downstream. The licence has to permit that.
Territory. Worldwide, or you have created a geographic restriction on a product that will be accessed globally.
Consent is a separate question from licence
A vendor can hold a perfectly valid commercial licence over data it had no right to collect. These are two different chains and both have to hold.
For personal data — and speech, faces and handwriting are all personal data — the individual must have consented to this specific use. That means consent naming machine-learning training, given in a language the person understands, at a reading level they can actually parse. We publish our consent standard in full for exactly this reason.
What provenance documentation should include
A licence is a promise. Provenance is the evidence. Ask any vendor for: collection date and location per record, a contributor pseudonym linking to a consent artefact, the consent version that governed that record, the compensation basis, and an unbroken chain of custody.
If a vendor cannot trace a single row back to a signature, they are asking you to accept their word. Some buyers will. Frontier labs and regulated enterprises will not.
Indemnity: what it does and does not cover
IP indemnity shifts some risk from you to the vendor, capped at a negotiated amount. It is worth having and it is not a substitute for diligence.
Watch the exclusions. Indemnity typically does not cover data you supplied, third-party licensed content where upstream terms govern, or use outside the licensed scope. A vendor who points those exclusions out proactively is more trustworthy than one who buries them in a schedule.
Answers
Frequently asked questions
A licence that explicitly permits using data to train, fine-tune, evaluate and distribute machine-learning models, including derivative models and including commercial use. It differs from a standard content licence, which typically grants reproduction and display rights drafted before model training existed as a use case.
It is legally contested and jurisdiction-dependent, and the position is actively shifting. The practical issue for a company is different from the theoretical one: scraped data usually has no consent chain and no provenance record, so it becomes a problem at the point of enterprise legal review, customer due diligence or acquisition — not at collection. This is not legal advice; ask your counsel.
An honest vendor will remove the records from the active corpus, stop further licensing, and notify affected clients — while stating plainly that a model already trained on the data cannot be un-trained. Be sceptical of any vendor implying otherwise; the promise is not deliverable.
Send us your counsel's redlines.
We would rather resolve licence questions before you scope a programme than after. The MSA and DPA go out the same day.
Average first response: under 6 business hours. NDAs signed same day.