New: conversational Urdu and Punjabi speech corpora now licensable  See the catalog →

Image & video data

Visual data from the streets, homes and roads of Pakistan

Faces, gestures, documents, retail, driving and liveness data captured in real Pakistani environments — with consented biometrics, in-region legal basis, and annotation to your schema.

Face & livenessGesture & poseDriving / ADASDocuments & OCRRetailMedical

Capture programs

Collected on location, not sourced from stock

Stock imagery of Pakistan is a photographer's idea of Pakistan. Model performance depends on what the camera will actually see at inference.

Face, liveness & anti-spoofing

Consented multi-pose, multi-lighting facial capture across ethnicities and age bands, plus print, replay, mask and deepfake attack sets for presentation-attack detection.

Gesture, pose & activity

Culturally specific gestures, sign languages, daily-activity video, and full-body pose sequences captured indoors and outdoors with synchronised multi-angle rigs.

Driving, ADAS & mobility

Instrumented-vehicle capture in mixed traffic: two-wheelers, auto-rickshaws, unmarked lanes, monsoon, night and informal roadside activity, with sync'd GPS and optional LiDAR.

Documents & OCR

Handwritten and printed forms, receipts, prescriptions, CNICs, utility bills and hand-kept ledgers in Nastaliq Urdu, Sindhi and Shahmukhi Punjabi, with reading-order annotation.

Retail, logistics & industrial

Shelf and planogram imagery from modern trade and informal kirana stores, warehouse and factory-floor footage, defect capture and packaging variation.

Medical & regulated

De-identified clinical imagery and procedural video collected under ethics-board approval with clinician oversight, in partnership with in-region institutions.

Annotation

Labeling that survives an audit

We annotate what we collect and what you already own. Same standard either way: written schema, calibration round, double-blind pass, adjudication, reported agreement.

  • 2D bounding boxes, rotated boxes and polygons
  • Semantic and instance segmentation at pixel level
  • Keypoints, landmarks and full-body pose
  • 3D cuboids and point-cloud annotation on LiDAR
  • Multi-object tracking with persistent IDs across frames
  • OCR transcription with reading order and script tagging
  • Native-language captioning and visual question answering
  • Attribute and hierarchical taxonomy labeling

See annotation & evaluation services

Example: liveness & PAD set

Typical delivery specification — all fields are configurable.

Subjects6,000 consented
Ethnic groups6 Pakistani groups
Age bands18–29 / 30–49 / 50+
Poses per subject12
Lighting conditions6 (incl. backlit, low-lux)
Devices18 handset models
Attack typesPrint, replay, 2D mask, 3D mask
Attack:bona-fide ratio1 : 3
Resolution≥ 1080p, EXIF retained
ConsentBiometric-specific, witnessed
Delivery6 weeks, 2 batches

Applications

What teams build with it

Use caseWhat we collectTypical scale
Face recognition & KYCMulti-pose consented facial capture across 9+ ethnic groups5k–100k subjects
Liveness / anti-spoofingBona-fide + print, replay, mask and deepfake attacks3k–20k subjects
ADAS & autonomous drivingInstrumented-vehicle video, mixed traffic, monsoon & night200–5,000 hrs
Document AI / OCRHandwritten and printed forms across 12+ scripts50k–2M pages
Retail shelf intelligenceModern-trade and informal store shelf imagery100k–3M images
Gesture & sign languageMulti-angle sequences with native signers50k–500k clips
Vision-language modelsImages with native-language captions and VQA pairs100k–5M pairs

Answers

Image & video questions

Facial and biometric collection runs under a separate, heightened consent that names biometric processing explicitly, states retention period and revocation rights, and is executed in the contributor's own language with a witnessed signature. We do not collect facial data in jurisdictions where we cannot secure a defensible legal basis, and we will tell you which countries those are before you plan a program around them.

Yes — that is most of what makes Pakistani visual data hard to source. We have fielded in sabzi mandis, kiryana stores, on rickshaws and motorcycles, in cotton and wheat fields, textile factories, hospital outpatient queues, madrasas and monsoon street conditions. Environment access is arranged through local partners with site permissions documented before capture.

Bounding boxes, polygons, semantic and instance segmentation, keypoints and pose, 3D cuboids on LiDAR, object tracking across frames, OCR transcription with reading order, and free-form captioning in the local language. Annotation can be applied to data we collect or to data you already own.

Yes, including the conditions that make Pakistani roads distinctive: rickshaws, motorcycles carrying whole families, unmarked lanes, decorated trucks, informal roadside commerce, livestock, monsoon and night-time low-light. We field instrumented vehicles with synchronised camera, GPS and optional LiDAR, and handle face and plate redaction before delivery.

An automated PII and sensitive-content pass runs on every frame before delivery: face detection, licence plate, ID document, screen content and address signage. Anything flagged is redacted or removed per your instruction. Annotators working on sensitive projects operate in secure facilities with no personal devices, and access is logged per asset.

Show us one frame your model gets wrong.

Send a failure case and we will design the capture program that covers it — environments, devices, subjects and volumes, with a realistic price.

Average first response: under 6 business hours. NDAs signed same day.

Free samples Get a quote