Speech & audio data · ready-made dataset
Wake word and command dataset
Wake-word models fail in two ways: they miss the trigger at distance and in noise, or they fire on something that merely sounds like it. This set is built around both failure modes — positive triggers across four distances and noise grades, and hard negatives chosen to be confusable.
Datasheet summary
What is in Wake Word & Command Set
Positive wake triggers at 0.3/1/3/5 m with graded noise, plus hard negatives and confusable phrases.
| Dataset ID | TDA-SP-005 |
|---|---|
| Modality | Speech |
| Languages / coverage | Urdu, Punjabi, Pashto, English |
| Typical first delivery | 50k–200k utt |
| Free sample | 2–5 hours of audio with transcripts and full metadata, within one business day |
| Provenance | Collected in Pakistan; collection date, location, contributor pseudonym, device profile and consent version on every record |
| Licence | Perpetual, worldwide, sublicensable; derivative-model rights stated explicitly. Exclusivity optional. Licensing detail |
| Delivery | Your S3, GCS or Azure bucket with manifests, checksums, QA report and datasheet |
| Guarantee | 14-day acceptance window; anything outside signed tolerance re-collected free |
Included in every delivery
- Positive wake triggers at 0.3 m, 1 m, 3 m and 5 m
- Graded background-noise conditions
- Hard negatives and confusable phrases
- Urdu, Punjabi, Pashto and English
Built for
- On-device wake-word detection for Pakistani-market devices
- Command-and-control for in-car, appliance and smart-speaker products
- False-accept reduction using confusable-phrase negatives
- Far-field robustness evaluation
Evaluate before you buy
Start with the free sample, then a pilot
Request 2–5 hours of audio with transcripts and full metadata from Wake Word & Command Set — identical in format and quality to the paid corpus, no watermark, no sales call. Pilots are creditable against production and start within five business days of a signed SOW.
Average first response under 6 business hours. NDA signed same day.
- Consent form, licence text and DPA available before you commit
- Datasheet lists composition, QA method and known limitations
- Subsets by language, region or condition on request
- Custom collection to your spec using the same protocol
Answers
Questions about this dataset
Yes. The catalog set covers a standard prompt list; a custom wake word or command grammar is a short custom collection using the same protocol and recruitment network.
Phrases chosen to be acoustically close to the trigger — near-homophones, partial matches and common speech that shares its onset — so the model learns to reject them rather than only learning to accept the positive.
Each utterance is captured at a fixed distance of 0.3, 1, 3 or 5 metres, and the distance is recorded in the metadata so you can train and evaluate per condition.
Related datasets
Also in the catalog
Conversational Urdu Speech
Urdu (Karachi, Lahori, standard)
SpeechPakistani Telephony Corpus
Urdu, Punjabi, Sindhi, Pashto
SpeechPakistani-Accented English
English, L1-tagged across 7 languages
See the full catalog · Speech & audio data services · Languages we cover