New: conversational Urdu and Punjabi speech corpora now licensable  See the catalog →
The Dataa · Collection & licensing Origin PK · Rev 2026.08
Collecting now · Datasheet 001

You need Asian data for AI model training.

Your model learned the West. It has never heard the 255 million people who speak Urdu, Punjabi, Pashto, Sindhi, Saraiki and Balochi — so we collect it at source, consent it at the record level, and license it for training.

Free samples in 1 business day · NDA signed same day · No sales gate

Provenance Pakistan Labelled on every record. No re-badging.
Speaker base 255M 5th largest country on earth.
Modalities Speech · Text
Image · Video
Plus annotation on data you own.
Consent Per record, versioned In-language, withdrawable, auditable.
Licence Perpetual · sublicensable Derivative-model rights stated explicitly.

Method

Nothing reaches you ungated

Four capture streams, three gates, one delivery. Every stage produces an artefact you receive.

Read the methodology

COLLECT VERIFY DELIVER Consent verifiedDouble-blind QAPII redacted SpeechTextImageVideo SIGNED DELIVERY with datasheet
Every stage below is contractual, not aspirational See the methodology

The operation

Six cities. One quality standard.

Every city runs its own recruitment, consent and QA staff, and everything they collect funnels through the same gates before it reaches you.

QA & DELIVERY KarachiUrdu, SindhiLahorePunjabiIslamabadPothwariPeshawarPashto, HindkoQuettaBalochi, BrahuiMultanSaraiki
Example collection records Sample records, not a live feed
SaraikiMultan · conversational · 48 minQA passed
PashtoPeshawar · telephony · 2-partyIn QA
UrduKarachi · Nastaliq handwriting · 340 pagesQA passed
BalochiQuetta · scripted prompts · 61 minRecording

What every delivery carries, from the first pilot batch onward

PER-RECORD CONSENT TRAINING-RIGHTS LICENCE DATASHEET INCLUDED NAMED QA THRESHOLDS FREE RE-COLLECTION

Quick answers

The questions buyers actually ask

Short, quotable answers. Free to cite — attribution to The Dataa appreciated.

What is Asian AI training data?

Asian AI training data is speech, text, image and video collected from Asian language speakers and licensed for training machine-learning models. It matters because mainstream corpora are overwhelmingly English and Western, so models systematically underperform for the billions of people who are not.

Where can I buy Urdu, Punjabi or Pashto training data?

The Dataa collects and licenses it directly at source. Speech, text, image and annotation data across Urdu, Punjabi, Saraiki, Sindhi, Pashto, Balochi and Brahui, with per-record consent and a training-rights licence. Free unwatermarked samples arrive within one business day, with no sales call required.

How much data do I need to fine-tune a model?

Fine-tuning a multilingual ASR model on a new language shows meaningful gains from 200–500 hours and keeps improving to a few thousand. A single-speaker TTS voice needs 20–60 studio hours. Instruction fine-tuning changes behaviour usefully from around 10,000 well-authored pairs. Training from scratch is a different order of magnitude and is rarely the right call for a low-resource language.

What makes a training dataset legally safe to use?

Three things: informed consent from each contributor that explicitly names machine-learning training, a licence granting perpetual and sublicensable rights including derivative models, and documented provenance you can audit back to a single record. Scraped data usually fails all three, which is why it becomes a problem at the point of enterprise legal review rather than at collection.

Why do models fail on low-resource languages?

Not architecture — data. Architectures and scaling laws transfer across languages; what does not transfer is data that was never collected. There is no Common Crawl of spoken Sindhi and no archive of how a Pashto speaker asks a bank for help, so the model has nothing to learn from.

Is commissioned data better than open datasets?

Not automatically. Use Common Voice, FLEURS and OpenSLR first — they are free and often sufficient. Commissioned collection is what you buy when the open data has run out, when you need a demographic distribution matching your actual users, or when the licence has to survive enterprise legal review.

The data gap

The fifth-largest country on earth, missing from your corpus

More people than Brazil. More than Russia and Japan combined. And its share of the corpora your model learned from rounds to zero — which is why the languages below are where the remaining headroom is.

We exist for the last mile: the Saraiki caller your IVR misroutes, the handwritten Urdu prescription your OCR cannot read, the Punjabi insult your safety classifier scores as neutral.

See the languages we cover

Where models quietly fail

FailureWhat the team seesWhat it actually is
Voice agent drops calls“Bad audio quality”ASR trained on studio speech, deployed to budget handsets in Karachi
OCR misreads forms“Poor scans”No training data for handwritten Nastaliq Urdu or Shahmukhi Punjabi
Safety classifier misses abuse“Low-resource edge case”Slurs are code-switched Urdu-English; no labeled examples exist
Assistant sounds foreign“Tone problem”SFT data translated from English, not authored by natives

Every one of these is a data problem before it is a modeling problem.

Any modality, any Pakistani language

If a person in Pakistan can produce it, we can collect it

Four practices, one operating standard: written spec, quota-controlled recruitment, double-blind QA, and a licence your counsel will actually sign.

Speech & Audio

Scripted prompts, spontaneous conversation, call-center recordings, wake words, far-field and in-vehicle audio, Urdu-English code-switching, and Pakistani-accented English.

Explore speech data

Text & NLP

Native-authored Urdu and regional-language corpora, Urdu–English translation pairs, instruction and SFT sets, RLHF rankings, and domain text in law, medicine and finance.

Explore text data

Image & Video

Faces with consented biometrics, gesture and pose, street and traffic footage, retail and shop scenes, handwritten Nastaliq documents, and liveness capture.

Explore visual data

Annotation & Evaluation

Transcription, translation, bounding boxes, segmentation, entity and intent labels, red-teaming, side-by-side model evals and human preference judgments.

Explore annotation

Multimodal & Agentic

Screen recordings, app-navigation traces, tool-use trajectories, document-plus-image QA pairs, and video with aligned Urdu narration.

Discuss a custom build

Off-the-Shelf Catalog

Pre-collected, pre-licensed corpora you can evaluate this week — with free unwatermarked sample packs and no procurement cycle.

Browse the catalog

Reach beyond the border

Pakistani data does not stop at Pakistan

We collect in Pakistan and we label every record that way, because provenance is the first thing a serious buyer checks. But a large share of what we collect carries real, measurable transfer value into neighbouring populations — and some of it barely degrades at all.

Ask us for the transfer study. We will share the WER deltas we have measured training Hindi ASR on Urdu data, and the equivalent for Punjabi across the border — including the cases where the gain was too small to justify the spend.

Where it carries over

The border is a line on a map

  • Punjabi is spoken on both sides of the Wagah border by roughly 150 million people. Shahmukhi and Gurmukhi differ in script, not in speech — audio transfers almost intact.
  • Urdu and Hindi are largely mutually intelligible in spoken form. Conversational Urdu ASR data measurably improves Hindi ASR, and the reverse.
  • Pashto and Dari span the Afghan border with no meaningful linguistic discontinuity.
  • Faces and skin tones across Pakistan and north India overlap heavily — a real advantage for face, liveness and medical vision models.
  • Acoustic environments — mixed traffic, open-air markets, budget Android handsets on weak networks — are near-identical across the region.

Start where you are

Three ways teams buy from us

Pick the path that matches your stage. Each one has a different entry point, a different contract, and a different first week.

Frontier labs

Pretraining-scale, licence-clean

Ten-thousand-hour speech programs, multi-year text licensing, and RLHF panels in languages nobody else staffs. Provenance documentation built for a legal review, not a slide.

  • Dedicated delivery pod & named program lead
  • Per-record consent artefacts and chain of custody
  • IP indemnity and exclusivity options
  • Monthly recurring collection commitments

Request a licensing review

Enterprise AI teams

Domain data, inside your perimeter

You have the use case and the compliance officer. We bring vetted collectors, secure facilities, and SLAs your procurement team has seen before.

  • MSA, DPA and security questionnaire ready to go
  • On-prem / VPC annotation for regulated data
  • Named SLAs on accuracy, throughput and turnaround
  • Vendor onboarding support & insurance certificates

Start vendor onboarding

Startups & research

Small, fast, priced up front

You need 200 hours and an answer today, not a six-week procurement cycle. Fixed unit prices, catalog data on a click-through licence, pilots that start next Monday.

  • Free unwatermarked sample packs, no call required
  • Small pilots, creditable against production
  • Ready-made datasets you can licence this week
  • Academic and pre-seed terms available

Start with free samples

How it works

Five stages. Every one of them auditable.

The difference between usable data and expensive noise is process discipline. Ours is written down, and you see the artefacts at each gate.

Read the full methodology

Specification

We co-write a data spec: demographics, dialect quotas, acoustic or visual conditions, label schema, edge cases, and explicit acceptance tolerances.

Week 0Signed spec doc

Recruitment & consent

Contributors are sourced through regional partners, ID-verified, screened against quota, and consented in their own language before recording a single second.

Week 1Consent artefacts

Collection

Field teams, studios or our mobile capture app — whichever the spec demands. Live dashboards show quota fill by language, gender, age band and region.

Weeks 2–nLive dashboard

Double-blind QA

Two independent annotators, adjudication on disagreement, plus automated checks for SNR, clipping, duplicates, PII leakage and label drift.

ContinuousIAA reported

Delivery & acceptance

Data lands in your bucket with manifests, checksums, a QA report and the licence. You get 14 days to accept; misses are re-collected free.

On schedule14-day acceptance

Honest comparison

Where the alternatives break down

You have three other options. Here is what each actually gets you when the target language is Saraiki, Balochi or Hindko.

Scrape the webGlobal crowd platformGeneric BPOThe Dataa
Low-resource language depthThinThinVariableDeep
Demographic quota controlNoneNoneLimitedEnforced
Training-rights licenceUnclearPlatform ToSClient-draftedExplicit, per record
Per-record consent artefactNoNoRarelyYes
Native-speaker QANoCrowd-gradedSometimesDouble-blind + adjudication
Re-collection guaranteen/aNoNegotiatedStandard in MSA
Time to first usable batchDays (unusable)2–4 weeks6–10 weeks2–3 weeks

Comparison reflects our experience in Pakistan's low-resource languages; results differ for high-resource English tasks, where a crowd platform is often the cheaper correct answer.

Before you commit

Don't trust us. Verify us.

We are a new supplier and we are not going to pretend otherwise. So the burden of proof sits with us, and every claim on this site is designed to be checked before you spend anything.

Start with a free sample

Unwatermarked sample data in your target language, delivered within one business day. No call, no NDA, no procurement process. Listen to it, run it through your pipeline, then decide whether to keep talking to us.

Request a sample

Read the consent pack first

The actual contributor consent form, the licence text and our DPA, sent before you commit to anything. Give them to your legal team early — that is the stage where most data purchases quietly die.

See the consent framework

Buy a pilot, not a corpus

A small, deliberately diverse pilot batch scoped against your full specification. Fine-tune, measure, and see which speaker segments actually moved. Pilot fees are credited against production work.

How a pilot runs

Hold us to written thresholds

Per-task accuracy thresholds, a 14-day acceptance window, and re-collection at our cost for anything outside signed tolerance. If we miss, it is our problem and our invoice.

See the guarantees

Trust & compliance

Collected ethically, or it does not ship

Fair pay is not a CSR line item for us — underpaid contributors produce rushed, low-variance data. Doing this properly is also what makes the data good.

  • Contributors paid at or above the local living wage, disclosed in writing before work begins
  • Plain-language consent in the contributor's own language, with a revocation channel
  • GDPR-ready processing, Pakistan's PECA and draft PDP Bill obligations, and EU standard contractual clauses
  • ISO 27001-aligned controls; encryption in transit and at rest; access on need-to-know
  • PII detection and redaction pass before any delivery leaves our environment
  • IP indemnity available on production contracts

Read our trust commitments

Provenance you can audit

Every record carries collection date, location, contributor pseudonym, device profile and consent version.

Licence written for training

Perpetual, worldwide, sublicensable, with derivative-model rights stated explicitly. Exclusivity optional.

Security that survives review

SSO, RBAC, audit logs, secure-facility options, and a completed CAIQ available under NDA.

Answers

Questions we get in the first call

If yours is not here, ask us directly — we answer in plain language.

Crowd platforms give you whoever signs up — which in Pakistan skews heavily toward urban, English-literate, 18–30 male users in Karachi and Lahore with good bandwidth. That is exactly the distribution that makes models fail on real users. We run quota-controlled field collection: we recruit to a demographic and acoustic spec you approve, in the districts where the speakers actually live, and we reject submissions that fall outside quota even when they are technically clean.

Yes, and that is the part we spend the most money on. Every contributor signs a plain-language consent in their own language that explicitly grants perpetual, worldwide, sublicensable rights for machine-learning training and derivative model distribution. Consent artefacts are stored per record and delivered with the dataset, so your legal team can audit any single row back to its signature.

A paid pilot: typically 50–200 hours of audio, 25k–100k text items, or 10k–50k images in one or two languages. Pilots run 3–5 weeks and are creditable against a production order. Free sample packs (no cost, no watermark, real data) are available for every catalog dataset, with no call required.

Off-the-shelf catalog data: same week, after a signed licence. Custom collection: kickoff within 5 business days of a signed SOW; first QA-passed batch usually lands in week 2–3 so you can validate the spec before we scale headcount.

Yes. We support delivery to your S3/GCS/Azure bucket, work-in-your-VPC arrangements, on-prem annotation for regulated clients, and secure facility work where no annotator has internet or personal devices on the floor. See Security.

Every delivery ships with an acceptance report against the spec you signed. You get a 14-day acceptance window; anything outside tolerance is re-collected or re-annotated free of charge, and we do not invoice for rejected units. That guarantee is written into our standard MSA, not just our marketing.

Tell us the language. We will tell you what it takes.

Send a two-line description of what you need. You will get a scoped approach and a realistic timeline back — not a discovery-call funnel.

Average first response: under 6 business hours. NDAs signed same day.

Free samples Get a quote