Text & NLP data
Text written by natives, not translated by machines
Original corpora, parallel translation pairs, instruction and preference data, and domain text in Pakistan's languages — authored under contract by paid native speakers with the training rights assigned to you.
What we produce
Six kinds of text, all natively authored
Monolingual corpora
Long-form original writing across registers — news, essay, forum, spoken transcript, technical, colloquial — balanced by topic and region for pretraining and continued pretraining.
Parallel & translation
Human translation pairs at document and sentence level, with back-translation QA, terminology glossaries and domain-specific pairs for legal, medical and financial text.
Instruction & SFT
Natively authored prompt/response pairs, multi-turn dialogue, tool-use and function-calling traces, and refusal examples grounded in local norms rather than translated ones.
RLHF & preference
Side-by-side rankings, rubric-scored responses, critique-and-revise chains, and constitutional-style preference data collected by calibrated native raters.
Safety & red-team
Culturally specific toxicity, slur and hate-speech lexicons, jailbreak attempts in-language, misinformation patterns, and region-specific harm taxonomies.
Reasoning & domain
Step-by-step solution traces, exam-style problems from national curricula, and expert-authored content in medicine, law, agriculture and finance.
The translationese problem
Translated SFT data teaches your model to be a tourist
A model fine-tuned on translated English instructions answers Indonesian questions with American assumptions: the wrong holidays, the wrong legal system, the wrong food, the wrong politeness register.
Native authoring costs more per item. It is also the only thing that produces a model users in the region describe as sounding “normal” rather than “foreign but fluent”.
Same question, two data sources
Illustrative comparison from a client evaluation.
Translated seed set
“To dispute a charge, contact your bank’s fraud department and file a chargeback under Regulation E within 60 days.”
Natively authored
“Report it in the app first so you get a ticket number, then call the branch — for UPI disputes the bank has to respond within the RBI turnaround, and you will need the transaction reference.”
Both are fluent. Only one is usable advice in the market you are shipping to.
Writer network
Who actually writes your data
Not an anonymous crowd. A contracted, tested, tiered pool of writers and reviewers with verified credentials where the domain requires them.
| Tier | Screening | Typical use | Review |
|---|---|---|---|
| Generalist writer | Native-speaker test, writing sample, plagiarism & AI-text screen | Corpora, conversational SFT, forum-register text | 1 reviewer |
| Senior writer | Above + editorial experience, 3-month track record, calibration round | Multi-turn dialogue, complex instruction, style-controlled text | 1 reviewer + spot audit |
| Domain expert | Verified licence or degree (MD, LLB, CA, MSc), CV check | Medical, legal, financial, agricultural and scientific content | 2 reviewers, 1 domain-qualified |
| Linguist / adjudicator | Applied-linguistics background, orthography specialism | Standard-setting, disagreement adjudication, QA rubric design | Sets the rubric |
| Safety rater | Enhanced screening, welfare briefing, rest requirements, opt-out at will | Toxicity, red-team, harm taxonomy, jailbreak collection | Blind double-rating |
Answers
Text & NLP questions
Written. Scraped low-resource text is a trap: it is dominated by machine translation, religious boilerplate and mirror sites, and it carries no licence you can rely on. Our corpora are authored or curated by paid native speakers under contract, with the training rights assigned to you in writing. Where we do license existing text, it comes from identified rights-holders with a signed agreement we can show you.
Yes, and this is our most-requested text service. We staff native-speaker writers who are domain-qualified — practising doctors for medical SFT, licensed lawyers for legal — and run every item through a second reviewer. Critically, prompts are authored natively rather than translated from an English seed set, so you get culturally grounded tasks instead of English tasks wearing a costume.
Native raters, calibrated against a written rubric in their own language, with a mandatory qualification round and ongoing gold-question injection. We report per-rater agreement and drop raters who drift. For safety-adjacent ranking we use a separate, more heavily screened and better-paid pool with rest requirements and opt-out at any point.
Constantly. Saraiki, Balochi, Brahui, Hindko and Shina all have contested or competing orthographies, and Punjabi is written in Shahmukhi here and Gurmukhi across the border. We agree a documented house standard with you before work starts, publish it to annotators, and enforce it in QA. Where it is useful we can deliver parallel transliterations in Latin script alongside the native script.
We guarantee it contractually and we test for it. Writers work in a monitored environment with paste-detection and keystroke telemetry, output is screened with multiple AI-text detectors plus stylometric checks, and a native reviewer flags translationese. Any writer found submitting model output is removed and their entire history is re-collected at our cost.
Give us one prompt in your target language.
We will return three natively authored responses and the translated equivalent, side by side, free. It is the fastest way to see the difference we are describing.
Average first response: under 6 business hours. NDAs signed same day.