Practical guides
What makes SFT data good, and why translated data is not
What separates useful instruction fine-tuning data from filler, how many pairs you need, and why translating an English seed set produces a model that sounds foreign.
Key takeaways
- Useful behaviour change typically starts around 10,000 well-authored pairs, not hundreds of thousands.
- Translated instruction data teaches a model to answer local questions with foreign assumptions.
- Diversity of task type matters more than volume. A thousand near-duplicate prompts teach almost nothing.
- Domain-expert authoring is non-negotiable for medical, legal and financial content.
How many pairs you actually need
Instruction fine-tuning shows useful behaviour change from roughly 10,000 well-authored prompt-and-response pairs. That surprises teams expecting six figures — the influential early instruction-tuning work demonstrated substantial gains at surprisingly modest scale, and the pattern has held.
Quality and diversity dominate. Ten thousand genuinely varied, carefully written pairs outperform a hundred thousand near-duplicates generated from a handful of templates, and they cost less to produce than the volume figure suggests.
The translationese problem
The standard shortcut is to take an English instruction set and machine-translate it. It is fast, cheap, and it produces a model that answers in the target language while thinking in English.
Concretely: a model fine-tuned on translated data will explain a chargeback under US Regulation E to a user asking about a UPI dispute, reference the wrong holidays, assume the wrong legal system, and use a politeness register that reads as subtly wrong to native speakers. Every sentence is fluent. The advice is useless.
Native authoring costs more per item and is the only thing that produces a model users describe as sounding normal rather than foreign-but-fluent.
What good SFT data contains
Task diversity. Not just question-answering: summarisation, extraction, rewriting, classification, multi-turn dialogue, tool use, structured output, and refusals.
Difficulty range. Trivial prompts teach little. So do impossible ones. The useful mass sits in the middle.
Realistic phrasing. Prompts written the way users actually type — including typos, incomplete context, and the terse register real people use — rather than the polished prose an annotator produces when trying to sound professional.
Grounded refusals. Examples of the model correctly declining, with reasons appropriate to local norms rather than transplanted from another culture.
Who should write it
For general instruction data, contracted native-speaker writers who have passed a writing test and a plagiarism and AI-text screen. For domain content, people with verifiable credentials — practising doctors for medical, licensed lawyers for legal, qualified accountants for financial.
This matters more than it sounds. A confident, fluent, wrong medical answer in a fine-tuning set does not just fail to help; it actively teaches the model to produce confident wrong medical answers in that language.
Detecting model-generated submissions
Writers submitting model output as human writing is a real and growing problem across the industry, and it quietly destroys the value of the dataset — you paid for human diversity and received model outputs, which is the opposite.
Defences that work: monitored writing environments with paste detection, keystroke telemetry, multiple AI-text detectors combined with stylometric checks, and native reviewers flagging translationese. When a writer is caught, their entire history should be re-collected.
Related pages on this site
Answers
Frequently asked questions
Useful behaviour change typically begins around 10,000 well-authored, genuinely diverse pairs. Quality and task diversity matter far more than raw count — a hundred thousand near-duplicates generated from a few templates will underperform a carefully varied ten thousand.
You can, and it produces a measurably worse model than native authoring. Translated instruction data carries the source culture's assumptions: wrong legal frameworks, wrong institutions, wrong politeness register. The output is fluent and the substance is frequently inapplicable.
Text that is grammatically fluent in the target language but carries the structure, idiom and cultural assumptions of the source. Native speakers identify it immediately as sounding foreign even when they cannot articulate why, and models trained on it inherit the quality.
Give us one prompt in your target Asian language.
We will return three natively authored responses and the translated equivalent, side by side, free. It is the fastest way to see the difference.
Average first response: under 6 business hours. NDAs signed same day.