New: conversational Urdu and Punjabi speech corpora now licensable  See the catalog →

Speech & audio data · ready-made dataset

Wake word and command dataset

Wake-word models fail in two ways: they miss the trigger at distance and in noise, or they fire on something that merely sounds like it. This set is built around both failure modes — positive triggers across four distances and noise grades, and hard negatives chosen to be confusable.

Free unwatermarked samplePer-record consentTraining-rights licence14-day acceptance

Datasheet summary

What is in Wake Word & Command Set

Positive wake triggers at 0.3/1/3/5 m with graded noise, plus hard negatives and confusable phrases.

Dataset IDTDA-SP-005
ModalitySpeech
Languages / coverageUrdu, Punjabi, Pashto, English
Typical first delivery50k–200k utt
Free sample2–5 hours of audio with transcripts and full metadata, within one business day
ProvenanceCollected in Pakistan; collection date, location, contributor pseudonym, device profile and consent version on every record
LicencePerpetual, worldwide, sublicensable; derivative-model rights stated explicitly. Exclusivity optional. Licensing detail
DeliveryYour S3, GCS or Azure bucket with manifests, checksums, QA report and datasheet
Guarantee14-day acceptance window; anything outside signed tolerance re-collected free

Included in every delivery

  • Positive wake triggers at 0.3 m, 1 m, 3 m and 5 m
  • Graded background-noise conditions
  • Hard negatives and confusable phrases
  • Urdu, Punjabi, Pashto and English

Built for

  • On-device wake-word detection for Pakistani-market devices
  • Command-and-control for in-car, appliance and smart-speaker products
  • False-accept reduction using confusable-phrase negatives
  • Far-field robustness evaluation

Evaluate before you buy

Start with the free sample, then a pilot

Request 2–5 hours of audio with transcripts and full metadata from Wake Word & Command Set — identical in format and quality to the paid corpus, no watermark, no sales call. Pilots are creditable against production and start within five business days of a signed SOW.

Average first response under 6 business hours. NDA signed same day.

  • Consent form, licence text and DPA available before you commit
  • Datasheet lists composition, QA method and known limitations
  • Subsets by language, region or condition on request
  • Custom collection to your spec using the same protocol

Answers

Questions about this dataset

Yes. The catalog set covers a standard prompt list; a custom wake word or command grammar is a short custom collection using the same protocol and recruitment network.

Phrases chosen to be acoustically close to the trigger — near-homophones, partial matches and common speech that shares its onset — so the model learns to reject them rather than only learning to accept the positive.

Each utterance is captured at a fixed distance of 0.3, 1, 3 or 5 metres, and the distance is recorded in the metadata so you can train and evaluate per condition.

Free samples Get a quote