New: conversational Urdu and Punjabi speech corpora now licensable  See the catalog →

Legal & ethics

What informed consent for AI training data actually requires

Valid consent must name machine-learning training, be understood by the person giving it, and be withdrawable. How that works at low literacy and across languages.

Key takeaways

  • Consent must name machine-learning training explicitly. A general “research purposes” clause does not cover it.
  • Valid consent requires actual comprehension, which at low literacy means reading the document aloud with a witness.
  • Biometric data — faces, voices — requires separate heightened consent naming biometric processing directly.
  • Withdrawal must be possible, and vendors should state honestly that a trained model cannot be un-trained.

The three conditions

Consent is valid when it is informed — the person understands what will happen to their data; specific — it names this use rather than gesturing at a category; and freely given — refusing carries no penalty and is a realistic option.

Each condition fails routinely in practice. A form written at university reading level given to someone with limited literacy is not informed. A clause saying “for research and product improvement” is not specific. Consent obtained from someone who believes refusal costs them work is not freely given.

Naming the use

The document must say that the data will be used to train machine-learning systems, and it should say so in ordinary words rather than legal abstraction. Where the resulting models may be commercial or distributed to third parties, the consent should say that too.

This is the clause most often missing from historical datasets, and it is why a great deal of legacy data cannot be relicensed for AI training even when the original collection was entirely legitimate for its stated purpose.

Consent at low literacy

This is the operationally hardest part and it is where most published ethics policies go quiet. If a contributor cannot read the form, a signature on it proves nothing.

Our approach: the consent is read aloud in the contributor's own language by a staff member who is not the person recruiting them, with an independent witness present, and the reading itself is recorded. Comprehension is then checked with three plain questions before any signature or thumbprint is taken. It is slower and more expensive than a checkbox, and it is the only version that would survive scrutiny.

Biometric data needs its own consent

Faces and voices are biometric identifiers and attract heightened protection under GDPR Article 9 and comparable regimes. A general data-collection consent does not cover them.

Biometric collection requires a separate document naming biometric processing explicitly, stating the retention period, and setting out revocation rights. Where a jurisdiction offers no defensible legal basis for the collection, the correct answer is to decline the work rather than proceed and hope.

Withdrawal, and being honest about its limits

Contributors must be able to withdraw, through a channel published in a language they read, without having to explain themselves.

On withdrawal, records should be removed from the active corpus, further licensing stopped, and downstream clients notified. And the consent document should state plainly what withdrawal cannot do: a model that has already ingested the data cannot be un-trained. Promising otherwise is a promise nobody can keep, and stating the limit up front is more respectful than implying an impossible guarantee.

Why this is also a quality argument

Ethics and quality are not in tension here. Underpaid, rushed, poorly-informed contributors produce low-variance, high-fraud data that is expensive to QA and frequently has to be re-collected.

Paying properly and consenting properly is not a cost centre offsetting quality — it is the cheapest quality intervention available. Every shortcut on consent or pay resurfaces later as fraud, homogeneity and rework.

Related pages on this site

Answers

Frequently asked questions

It applies where EU data subjects are involved, or where an EU-established controller or processor is in the chain, and it governs transfers into the EU. Many vendors — including us — apply GDPR as the working standard globally because it is the strictest regime routinely encountered, so meeting it satisfies the others.

A guardian can give consent, but best practice requires the child's own assent alongside it, plus additional safeguarding review, a chaperone present throughout, and shorter session caps. Many projects specifying child speech do not actually require it, and a good vendor will ask whether it is genuinely necessary.

Ask to see the actual consent template, in the original language. Ask how comprehension is verified at low literacy. Ask whether consent artefacts are stored per record and delivered with the dataset. Ask what the withdrawal channel is and what happens when someone uses it. A vendor who cannot answer all four concretely is asking you to take their word for it.

Ask to see a consent form before you buy.

From us or from anyone. It is the single most revealing question you can put to a data vendor.

Average first response: under 6 business hours. NDAs signed same day.

Free samples Get a quote