New: conversational Urdu and Punjabi speech corpora now licensable  See the catalog →

Image & video data · ready-made dataset

Pakistani face and liveness set

Face and liveness models trained on Western and East Asian datasets underperform on South Asian faces, and the consent behind many biometric datasets does not survive a legal review. This set is captured with explicit biometric consent from subjects across Pakistan's major ethnic groups, with presentation attacks recorded on the same subjects.

Free unwatermarked samplePer-record consentTraining-rights licence14-day acceptance

Datasheet summary

What is in Pakistani Face & Liveness Set

Consented multi-pose, multi-lighting facial capture with print, replay, 2D and 3D mask presentation attacks.

Dataset IDTDA-IV-001
ModalityImage
Languages / coveragePunjabi, Sindhi, Pashtun, Baloch, Muhajir
Typical first delivery500–5,000 subjects
Free sample500 images with annotations and full metadata, within one business day
ProvenanceCollected in Pakistan; collection date, location, contributor pseudonym, device profile and consent version on every record
LicencePerpetual, worldwide, sublicensable; derivative-model rights stated explicitly. Exclusivity optional. Licensing detail
DeliveryYour S3, GCS or Azure bucket with manifests, checksums, QA report and datasheet
Guarantee14-day acceptance window; anything outside signed tolerance re-collected free

Included in every delivery

  • Multi-pose, multi-lighting facial capture
  • Punjabi, Sindhi, Pashtun, Baloch and Muhajir subjects
  • Print, replay, 2D and 3D mask presentation attacks
  • Explicit per-subject biometric consent artefacts

Built for

  • Face recognition and verification for South Asian populations
  • Liveness and presentation-attack detection
  • Bias measurement across ethnic groups
  • KYC and identity-verification products in South Asia and the diaspora

Evaluate before you buy

Start with the free sample, then a pilot

Request 500 images with annotations and full metadata from Pakistani Face & Liveness Set — identical in format and quality to the paid corpus, no watermark, no sales call. Pilots are creditable against production and start within five business days of a signed SOW.

Average first response under 6 business hours. NDA signed same day.

  • Consent form, licence text and DPA available before you commit
  • Datasheet lists composition, QA method and known limitations
  • Subsets by language, region or condition on request
  • Custom collection to your spec using the same protocol

Answers

Questions about this dataset

Yes. Every subject signs a plain-language consent in their own language that names biometric use and machine-learning training, and the consent artefact is delivered per record.

Print, screen replay, 2D mask and 3D mask presentation attacks, captured on the same subjects as the genuine samples.

A typical first delivery covers 500–5,000 subjects; the datasheet gives the exact ethnic, age and gender composition.

Free samples Get a quote