Research & field notes
What we have learned collecting AI training data
20 practical write-ups on how much data a model actually needs, what a training licence has to say, why low-resource languages fail, and how collection is really done. No gated PDFs, no lead-capture walls.
Practical guides
Practical guides
How much audio do you need to fine-tune an ASR model?
How many hours of speech data you actually need to fine-tune an ASR model, with realistic thresholds for 200, 500, 2,000 and 10,000 hours.…
8 min read ·
Free datasets vs commissioned collection: when to pay
Common Voice, FLEURS and OpenSLR are free and often sufficient. Here is exactly when open data runs out and commissioned collection becomes …
8 min read ·
Quota sampling beats volume: how to specify a dataset
Why 500 well-distributed hours beat 5,000 convenience-sampled ones, which axes to quota on, and how to write a data specification that survi…
7 min read ·
How much data does a TTS voice need?
How many studio hours a text-to-speech voice actually needs, why TTS data differs completely from ASR data, and what phonetic balance means …
7 min read ·
What makes SFT data good, and why translated data is not
What separates useful instruction fine-tuning data from filler, how many pairs you need, and why translating an English seed set produces a …
8 min read ·
How to choose an AI training data vendor: 12 questions
Twelve questions that separate serious data vendors from resellers: provenance, consent, QA methodology, pay transparency and what happens w…
8 min read ·
How to run a speech data collection project
End to end: writing the spec, recruiting to quota, choosing capture conditions, transcription standards, QA gates and acceptance testing.…
9 min read ·
How to actually measure data annotation quality
Gold seeding, double-blind passes and senior adjudication — the three layers that produce trustworthy labels, and what a delivery report mus…
8 min read ·
Research
Research
Why AI models fail on low-resource languages
Low-resource languages are a data problem, not a modelling problem. What actually breaks, why speaker count and corpus size are unrelated, a…
7 min read ·
Why Urdu OCR is so much harder than Arabic OCR
Nastaliq breaks OCR pipelines built for Arabic. Sloping baselines, extreme ligature context and overlapping glyphs explain why, and what tra…
7 min read ·
Code-switching breaks speech recognition. Here is how to fix it
Bilingual speakers mix languages mid-sentence constantly. Why monolingual ASR fails on it, and how to specify and collect code-switched trai…
7 min read ·
RLHF outside English: why translated rubrics fail
Collecting preference data outside English needs calibrated native raters and a rubric written in their language. Why translated rubrics pro…
7 min read ·
Does Urdu data improve Hindi models? What transfers and what does not
Urdu and Hindi are largely mutually intelligible in speech and written in different scripts. What that means for ASR, OCR and LLM transfer, …
7 min read ·
Synthetic vs real training data: where synthetic actually works
Synthetic data is cheap and increasingly good. Where it genuinely substitutes for real collection, where it quietly degrades models, and how…
7 min read ·
Legal & ethics
Legal & ethics
AI training data licensing: what your contract must actually say
What a training-rights licence must contain, why most content licences do not cover model training, and the clauses that stop deals in enter…
9 min read ·
What informed consent for AI training data actually requires
Valid consent must name machine-learning training, be understood by the person giving it, and be withdrawable. How that works at low literac…
8 min read ·
Definitions
Definitions
Word Error Rate (WER) explained, and why yours is misleading
What WER measures, how to compute it, what counts as good, and why a model at 8% WER in testing can exceed 40% on real production audio.…
6 min read ·
Inter-annotator agreement: what the number actually tells you
Cohen's kappa, Krippendorff's alpha and raw agreement — what each measures, what counts as acceptable, and why low agreement usually means b…
6 min read ·
Language guides
Language guides
Punjabi NLP: 90 million speakers, almost no data
Punjabi is among the world's most spoken languages and among the least resourced in AI. What exists, what does not, and how Shahmukhi and Gu…
7 min read ·
Pashto NLP: dialects, script and the data that exists
Pashto spans Pakistan and Afghanistan with two major dialect groups that differ enough to degrade a model tuned on only one. What exists and…
7 min read ·
Have a question none of these answer?
Send it. We answer in plain language, and if the answer is useful to other people we will write it up.
Average first response: under 6 business hours. NDAs signed same day.