Off-the-shelf catalog
Dataset programmes ready to field
Standing collection programmes: the protocol, recruitment network, consent pack and QA design already exist for each one, so fielding starts in days rather than months. Free unwatermarked samples, no discovery call required.
Available now
Current catalog
Each entry ships with a full datasheet: collection window, demographic composition, QA methodology, known limitations and the licence text.
| Dataset | Type | Languages | Typical first delivery | |
|---|---|---|---|---|
| Conversational Urdu Speech TDA-SP-001 Two-party spontaneous telephone and in-person conversation, verbatim transcripts, speaker diarization, balanced across 4 age bands and urban/rural. | Speech | Urdu (Karachi, Lahori, standard) | 200–500 hrs | Request sample |
| Pakistani Telephony Corpus TDA-SP-002 8 kHz narrowband call-center audio, channel-separated agent/caller stems, GSM and VoIP codecs, intent and outcome labels. | Speech | Urdu, Punjabi, Sindhi, Pashto | 200–500 hrs | Request sample |
| Pakistani-Accented English TDA-SP-003 Read and spontaneous English graded by CEFR band, balanced by first language and region, with L1-tagged transcripts. | Speech | English, L1-tagged across 7 languages | 150–400 hrs | Request sample |
| Regional Languages Speech Pack TDA-SP-004 The least-resourced languages we field, collected in Multan, Quetta and Hazara with community-network recruitment. | Speech | Saraiki, Balochi, Brahui, Hindko | 100–300 hrs | Request sample |
| Wake Word & Command Set TDA-SP-005 Positive wake triggers at 0.3/1/3/5 m with graded noise, plus hard negatives and confusable phrases. | Speech | Urdu, Punjabi, Pashto, English | 50k–200k utt | Request sample |
| Native Urdu Instruction & SFT Corpus TDA-TX-001 Natively authored single- and multi-turn instruction pairs, culturally grounded, with refusal and safety examples. | Text | Urdu, plus 5 regional languages | 10k–100k items | Request sample |
| RLHF Preference Pairs TDA-TX-002 Side-by-side rankings by calibrated native raters, with rubric scores, rationales and per-rater agreement metadata. | Text | Urdu, Punjabi, Sindhi, Pashto | 5k–50k pairs | Request sample |
| Urdu–English Parallel Corpus TDA-TX-003 Human-translated document- and sentence-aligned pairs with terminology glossaries across law, medicine, finance, news, government and colloquial registers. | Text | Urdu↔English, 6 domains | 100k–1M pairs | Request sample |
| Pakistani Safety & Toxicity Lexicon TDA-TX-004 Culturally specific slurs, hate speech, harassment, sectarian language and coded terms, severity-graded by 3 native raters each. | Text | Urdu, Punjabi, Pashto, Sindhi | 10k–50k items | Request sample |
| Pakistani Face & Liveness Set TDA-IV-001 Consented multi-pose, multi-lighting facial capture with print, replay, 2D and 3D mask presentation attacks. | Image | Punjabi, Sindhi, Pashtun, Baloch, Muhajir | 500–5,000 subjects | Request sample |
| Nastaliq & Handwriting OCR TDA-IV-002 Printed Nastaliq and handwritten forms, receipts, prescriptions and ledgers with reading-order and field-level annotation. | Image | Urdu, Sindhi, Shahmukhi Punjabi | 10k–100k pages | Request sample |
| Pakistani Street & Traffic Footage TDA-IV-003 Instrumented-vehicle capture with rickshaws, motorcycles, unmarked lanes, monsoon and night, sync'd GPS, faces and plates redacted. | Video | Karachi, Lahore, Peshawar, Multan | 50–300 hrs | Request sample |
| Kiryana & Informal Retail Imagery TDA-IV-004 Kiryana store and roadside stall shelves with SKU-level boxes, occlusion and lighting variation. | Image | 6 Pakistani cities | 20k–200k images | Request sample |
Subset filtering is available on every dataset — filter by language, dialect, region, demographic band or device class and licence only the slice you need. Licensing terms are agreed per project; tell us the scope and we will come back with terms.
What ships with it
Every dataset comes with its own datasheet
Modelled on Datasheets for Datasets, because “12,000 hours of Indonesian” tells you almost nothing useful.
Composition
Speaker or subject counts, demographic breakdown across every axis, dialect distribution, device and environment mix.
Provenance
Collection window, locations, consent version, contributor compensation basis and chain of custody.
Quality
QA methodology, inter-annotator agreement, gold-set accuracy and the known error modes we found.
Limitations
An explicit section on what the dataset is not good for. We would rather you not buy it than misuse it.
Free sample pack
Real data. No watermark. No sales call.
Pick any catalog datasets and we will send a genuine slice — typically 2–5 hours of audio, 1,000 text items or 500 images — with the full metadata and datasheet, usually within one business day.
- Identical quality and format to the paid corpus
- Full metadata and datasheet included
- No watermarking or deliberate degradation
- Evaluation licence; no obligation to buy
Answers
Catalog questions
Yes. A sample pack is real data collected under the same protocol, consent framework and QA process as a full delivery — typically a few hours of audio, around a thousand text items or several hundred images. Same metadata, same format, no watermark, no degraded quality and no sales call required. It is deliberately small enough that you can evaluate it properly and we are not giving away a corpus.
Not by default. Catalog datasets are licensed non-exclusively to multiple clients, which is what makes them faster and cheaper to access than commissioning collection from scratch. If exclusivity matters to you, we can quote an exclusivity option on a catalog set, or you can commission custom collection where exclusivity is available from the start. Either way it is a conversation, not a price list.
Because each programme is fielded to your specification rather than pulled from a warehouse, what you receive is collected now, against current devices, current network conditions and current speech. Every delivery carries its collection window in the datasheet. Where we reuse material from an earlier collection we say so explicitly and give you its date rather than presenting it as fresh.
Yes. Filter by language, dialect, region, demographic band, device class, environment or label class, and pay only for the slice. Minimum order applies per dataset but is usually a small fraction of the whole.
Need a slice that is not in the catalog?
Subset filtering is free, and if the corpus you need does not exist yet, we will scope building it.
Average first response: under 6 business hours. NDAs signed same day.