Glossary
AI training data, defined in plain English
Every term a buyer, a lawyer or a journalist runs into when commissioning training data — written to be understood rather than to sound expert. Quote any of it; attribution appreciated.
AI training data
Data used to teach a machine-learning model. For language and speech models this means recorded audio, written text, labelled images or human judgements. Quality, licensing and demographic distribution matter more than raw volume, because a model can only learn the world its data describes.
Low-resource language
A language with little machine-readable data available, regardless of how many people speak it. Punjabi has roughly 90 million speakers in Pakistan and is still low-resource, because speaker count and digital corpus size are almost unrelated.
ASR (Automatic Speech Recognition)
Converting spoken audio into text. ASR quality is measured in Word Error Rate, and degrades sharply when deployment audio differs from training audio in accent, dialect, device, background noise or bandwidth.
WER (Word Error Rate)
The standard ASR accuracy metric: the proportion of words inserted, deleted or substituted relative to a reference transcript. Lower is better. A model at 8% WER on studio speech can easily exceed 40% on the same language over a phone line.
TTS (Text-to-Speech)
Generating spoken audio from written text. Needs clean, consistent, phonetically balanced studio recordings from a single speaker — typically 20 to 60 hours per voice.
SFT (Supervised Fine-Tuning)
Training a pretrained language model on curated prompt-and-response pairs so it follows instructions. Quality depends heavily on whether the pairs were natively authored or machine-translated from English.
RLHF (Reinforcement Learning from Human Feedback)
Aligning a model to human preferences by having people rank competing outputs, then training a reward model on those rankings. Outside English this requires calibrated native raters, not crowd workers using a translated rubric.
Preference data
Pairs or sets of model outputs ranked by humans for helpfulness, accuracy or safety. The raw material for RLHF and DPO.
Code-switching
Alternating between two languages within a single sentence or conversation — extremely common in Pakistani speech, where Urdu and English mix freely. Models trained on monolingual data typically fail on it.
Nastaliq
The calligraphic style of the Perso-Arabic script used to write Urdu. Its sloping, heavily context-dependent letterforms make it substantially harder for OCR than the Naskh style used for Arabic.
Quota sampling
Recruiting data contributors against a predefined demographic distribution rather than accepting whoever volunteers. Prevents the corpus from over-representing young, urban, male, high-bandwidth participants.
Convenience sampling
Collecting from whoever is easiest to reach. Cheap and fast, and the single most common cause of models that benchmark well but fail in production.
Inter-annotator agreement
How consistently independent annotators produce the same label on the same item, reported as Cohen's kappa or Krippendorff's alpha. Low agreement usually means the labelling guidelines are ambiguous, not that the annotators are careless.
Gold questions
Pre-adjudicated items secretly seeded into live annotation work to measure each annotator's accuracy continuously.
Double-blind QA
Having two annotators label the same item independently, with neither able to see the other's work, and routing disagreements to a senior adjudicator.
Informed consent
A contributor's documented agreement to their data being collected and used for a stated purpose — for AI data, that purpose must name machine-learning training explicitly. Valid consent requires the person to actually understand it, which at low literacy means reading it aloud with a witness.
Training-rights licence
A licence that explicitly permits using data to train, fine-tune and evaluate machine-learning models, including derivative models. Many content licences predate this use case and do not clearly grant it.
Provenance
The documented origin and chain of custody of a dataset — where it was collected, when, from whom, under what consent. The first thing a serious buyer's legal team checks.
Datasheet for Datasets
A standardised document describing a dataset's composition, collection process, intended uses and limitations. Proposed by Gebru et al. and now expected practice for responsibly published data.
PII redaction
Detecting and removing personally identifying information — names, ID numbers, faces, licence plates, addresses — before a dataset is delivered.
Far-field audio
Speech recorded at a distance from the microphone, typically one to five metres. Sounds very different from close-talk audio and requires its own training data.
Liveness / anti-spoofing data
Imagery used to teach a system to distinguish a real person from a photograph, screen replay or mask. Requires both genuine captures and deliberate attack samples.
Transfer value
How much a model trained on one language or population improves on a related one. Pakistani Punjabi audio transfers strongly to Indian Punjabi; Urdu text transfers poorly to Hindi OCR because the scripts differ.
Data cascade
The compounding downstream damage caused by upstream data problems — a labelling ambiguity in week one becomes a systematic model failure six months later.
Definitions written by The Dataa. Free to quote with attribution to The Dataa.
Still not sure what you need?
Describe the model behaviour you are trying to fix in one sentence. We will tell you which of these things applies and which do not.
Average first response: under 6 business hours. NDAs signed same day.