Image & video data
Visual data from the streets, homes and roads of Pakistan
Faces, gestures, documents, retail, driving and liveness data captured in real Pakistani environments — with consented biometrics, in-region legal basis, and annotation to your schema.
Capture programs
Collected on location, not sourced from stock
Stock imagery of Pakistan is a photographer's idea of Pakistan. Model performance depends on what the camera will actually see at inference.
Face, liveness & anti-spoofing
Consented multi-pose, multi-lighting facial capture across ethnicities and age bands, plus print, replay, mask and deepfake attack sets for presentation-attack detection.
Gesture, pose & activity
Culturally specific gestures, sign languages, daily-activity video, and full-body pose sequences captured indoors and outdoors with synchronised multi-angle rigs.
Driving, ADAS & mobility
Instrumented-vehicle capture in mixed traffic: two-wheelers, auto-rickshaws, unmarked lanes, monsoon, night and informal roadside activity, with sync'd GPS and optional LiDAR.
Documents & OCR
Handwritten and printed forms, receipts, prescriptions, CNICs, utility bills and hand-kept ledgers in Nastaliq Urdu, Sindhi and Shahmukhi Punjabi, with reading-order annotation.
Retail, logistics & industrial
Shelf and planogram imagery from modern trade and informal kirana stores, warehouse and factory-floor footage, defect capture and packaging variation.
Medical & regulated
De-identified clinical imagery and procedural video collected under ethics-board approval with clinician oversight, in partnership with in-region institutions.
Annotation
Labeling that survives an audit
We annotate what we collect and what you already own. Same standard either way: written schema, calibration round, double-blind pass, adjudication, reported agreement.
- 2D bounding boxes, rotated boxes and polygons
- Semantic and instance segmentation at pixel level
- Keypoints, landmarks and full-body pose
- 3D cuboids and point-cloud annotation on LiDAR
- Multi-object tracking with persistent IDs across frames
- OCR transcription with reading order and script tagging
- Native-language captioning and visual question answering
- Attribute and hierarchical taxonomy labeling
Example: liveness & PAD set
Typical delivery specification — all fields are configurable.
Applications
What teams build with it
| Use case | What we collect | Typical scale |
|---|---|---|
| Face recognition & KYC | Multi-pose consented facial capture across 9+ ethnic groups | 5k–100k subjects |
| Liveness / anti-spoofing | Bona-fide + print, replay, mask and deepfake attacks | 3k–20k subjects |
| ADAS & autonomous driving | Instrumented-vehicle video, mixed traffic, monsoon & night | 200–5,000 hrs |
| Document AI / OCR | Handwritten and printed forms across 12+ scripts | 50k–2M pages |
| Retail shelf intelligence | Modern-trade and informal store shelf imagery | 100k–3M images |
| Gesture & sign language | Multi-angle sequences with native signers | 50k–500k clips |
| Vision-language models | Images with native-language captions and VQA pairs | 100k–5M pairs |
Answers
Image & video questions
Facial and biometric collection runs under a separate, heightened consent that names biometric processing explicitly, states retention period and revocation rights, and is executed in the contributor's own language with a witnessed signature. We do not collect facial data in jurisdictions where we cannot secure a defensible legal basis, and we will tell you which countries those are before you plan a program around them.
Yes — that is most of what makes Pakistani visual data hard to source. We have fielded in sabzi mandis, kiryana stores, on rickshaws and motorcycles, in cotton and wheat fields, textile factories, hospital outpatient queues, madrasas and monsoon street conditions. Environment access is arranged through local partners with site permissions documented before capture.
Bounding boxes, polygons, semantic and instance segmentation, keypoints and pose, 3D cuboids on LiDAR, object tracking across frames, OCR transcription with reading order, and free-form captioning in the local language. Annotation can be applied to data we collect or to data you already own.
Yes, including the conditions that make Pakistani roads distinctive: rickshaws, motorcycles carrying whole families, unmarked lanes, decorated trucks, informal roadside commerce, livestock, monsoon and night-time low-light. We field instrumented vehicles with synchronised camera, GPS and optional LiDAR, and handle face and plate redaction before delivery.
An automated PII and sensitive-content pass runs on every frame before delivery: face detection, licence plate, ID document, screen content and address signage. Anything flagged is redacted or removed per your instruction. Annotators working on sensitive projects operate in secure facilities with no personal devices, and access is logged per asset.
Show us one frame your model gets wrong.
Send a failure case and we will design the capture program that covers it — environments, devices, subjects and volumes, with a realistic price.
Average first response: under 6 business hours. NDAs signed same day.