New: conversational Urdu and Punjabi speech corpora now licensable  See the catalog →

Off-the-shelf catalog

Dataset programmes ready to field

Standing collection programmes: the protocol, recruitment network, consent pack and QA design already exist for each one, so fielding starts in days rather than months. Free unwatermarked samples, no discovery call required.

Spec already writtenFree samplesFilter by subsetPilot before you scale

Available now

Current catalog

Each entry ships with a full datasheet: collection window, demographic composition, QA methodology, known limitations and the licence text.

DatasetTypeLanguages Typical first delivery
Conversational Urdu Speech
TDA-SP-001
Two-party spontaneous telephone and in-person conversation, verbatim transcripts, speaker diarization, balanced across 4 age bands and urban/rural.
SpeechUrdu (Karachi, Lahori, standard)200–500 hrsRequest sample
Pakistani Telephony Corpus
TDA-SP-002
8 kHz narrowband call-center audio, channel-separated agent/caller stems, GSM and VoIP codecs, intent and outcome labels.
SpeechUrdu, Punjabi, Sindhi, Pashto200–500 hrsRequest sample
Pakistani-Accented English
TDA-SP-003
Read and spontaneous English graded by CEFR band, balanced by first language and region, with L1-tagged transcripts.
SpeechEnglish, L1-tagged across 7 languages150–400 hrsRequest sample
Regional Languages Speech Pack
TDA-SP-004
The least-resourced languages we field, collected in Multan, Quetta and Hazara with community-network recruitment.
SpeechSaraiki, Balochi, Brahui, Hindko100–300 hrsRequest sample
Wake Word & Command Set
TDA-SP-005
Positive wake triggers at 0.3/1/3/5 m with graded noise, plus hard negatives and confusable phrases.
SpeechUrdu, Punjabi, Pashto, English50k–200k uttRequest sample
Native Urdu Instruction & SFT Corpus
TDA-TX-001
Natively authored single- and multi-turn instruction pairs, culturally grounded, with refusal and safety examples.
TextUrdu, plus 5 regional languages10k–100k itemsRequest sample
RLHF Preference Pairs
TDA-TX-002
Side-by-side rankings by calibrated native raters, with rubric scores, rationales and per-rater agreement metadata.
TextUrdu, Punjabi, Sindhi, Pashto5k–50k pairsRequest sample
Urdu–English Parallel Corpus
TDA-TX-003
Human-translated document- and sentence-aligned pairs with terminology glossaries across law, medicine, finance, news, government and colloquial registers.
TextUrdu↔English, 6 domains100k–1M pairsRequest sample
Pakistani Safety & Toxicity Lexicon
TDA-TX-004
Culturally specific slurs, hate speech, harassment, sectarian language and coded terms, severity-graded by 3 native raters each.
TextUrdu, Punjabi, Pashto, Sindhi10k–50k itemsRequest sample
Pakistani Face & Liveness Set
TDA-IV-001
Consented multi-pose, multi-lighting facial capture with print, replay, 2D and 3D mask presentation attacks.
ImagePunjabi, Sindhi, Pashtun, Baloch, Muhajir500–5,000 subjectsRequest sample
Nastaliq & Handwriting OCR
TDA-IV-002
Printed Nastaliq and handwritten forms, receipts, prescriptions and ledgers with reading-order and field-level annotation.
ImageUrdu, Sindhi, Shahmukhi Punjabi10k–100k pagesRequest sample
Pakistani Street & Traffic Footage
TDA-IV-003
Instrumented-vehicle capture with rickshaws, motorcycles, unmarked lanes, monsoon and night, sync'd GPS, faces and plates redacted.
VideoKarachi, Lahore, Peshawar, Multan50–300 hrsRequest sample
Kiryana & Informal Retail Imagery
TDA-IV-004
Kiryana store and roadside stall shelves with SKU-level boxes, occlusion and lighting variation.
Image6 Pakistani cities20k–200k imagesRequest sample

Subset filtering is available on every dataset — filter by language, dialect, region, demographic band or device class and licence only the slice you need. Licensing terms are agreed per project; tell us the scope and we will come back with terms.

What ships with it

Every dataset comes with its own datasheet

Modelled on Datasheets for Datasets, because “12,000 hours of Indonesian” tells you almost nothing useful.

Composition

Speaker or subject counts, demographic breakdown across every axis, dialect distribution, device and environment mix.

Provenance

Collection window, locations, consent version, contributor compensation basis and chain of custody.

Quality

QA methodology, inter-annotator agreement, gold-set accuracy and the known error modes we found.

Limitations

An explicit section on what the dataset is not good for. We would rather you not buy it than misuse it.

Free sample pack

Real data. No watermark. No sales call.

Pick any catalog datasets and we will send a genuine slice — typically 2–5 hours of audio, 1,000 text items or 500 images — with the full metadata and datasheet, usually within one business day.

  • Identical quality and format to the paid corpus
  • Full metadata and datasheet included
  • No watermarking or deliberate degradation
  • Evaluation licence; no obligation to buy

Typically delivered within one business day.

Request received. Your sample pack is on its way — check your inbox within one business day. If you need it faster, reply to the confirmation email and we will expedite.

Demo note: this form is a front-end demo. Connect it to your backend or a form service before launch.

Answers

Catalog questions

Yes. A sample pack is real data collected under the same protocol, consent framework and QA process as a full delivery — typically a few hours of audio, around a thousand text items or several hundred images. Same metadata, same format, no watermark, no degraded quality and no sales call required. It is deliberately small enough that you can evaluate it properly and we are not giving away a corpus.

Not by default. Catalog datasets are licensed non-exclusively to multiple clients, which is what makes them faster and cheaper to access than commissioning collection from scratch. If exclusivity matters to you, we can quote an exclusivity option on a catalog set, or you can commission custom collection where exclusivity is available from the start. Either way it is a conversation, not a price list.

Because each programme is fielded to your specification rather than pulled from a warehouse, what you receive is collected now, against current devices, current network conditions and current speech. Every delivery carries its collection window in the datasheet. Where we reuse material from an earlier collection we say so explicitly and give you its date rather than presenting it as fresh.

Yes. Filter by language, dialect, region, demographic band, device class, environment or label class, and pay only for the slice. Minimum order applies per dataset but is usually a small fraction of the whole.

Need a slice that is not in the catalog?

Subset filtering is free, and if the corpus you need does not exist yet, we will scope building it.

Average first response: under 6 business hours. NDAs signed same day.

Free samples Get a quote