Image & video data · ready-made dataset
Nastaliq and handwriting OCR dataset
Nastaliq breaks OCR pipelines built for Arabic: sloping baselines, extreme ligature context and overlapping glyphs. Handwritten Nastaliq is harder again. This set pairs real printed and handwritten documents with reading-order and field-level annotation so a document-AI model can learn the script as it is actually written.
Datasheet summary
What is in Nastaliq & Handwriting OCR
Printed Nastaliq and handwritten forms, receipts, prescriptions and ledgers with reading-order and field-level annotation.
| Dataset ID | TDA-IV-002 |
|---|---|
| Modality | Image |
| Languages / coverage | Urdu, Sindhi, Shahmukhi Punjabi |
| Typical first delivery | 10k–100k pages |
| Free sample | 500 images with annotations and full metadata, within one business day |
| Provenance | Collected in Pakistan; collection date, location, contributor pseudonym, device profile and consent version on every record |
| Licence | Perpetual, worldwide, sublicensable; derivative-model rights stated explicitly. Exclusivity optional. Licensing detail |
| Delivery | Your S3, GCS or Azure bucket with manifests, checksums, QA report and datasheet |
| Guarantee | 14-day acceptance window; anything outside signed tolerance re-collected free |
Included in every delivery
- Printed Nastaliq and handwritten documents
- Forms, receipts, prescriptions and ledgers
- Reading-order annotation
- Field-level annotation
- Urdu, Sindhi and Shahmukhi Punjabi
Built for
- Urdu, Sindhi and Shahmukhi Punjabi OCR
- Handwriting recognition for prescriptions, forms and ledgers
- Document-AI field extraction and layout understanding
- Digitisation of paper records in South Asia
Evaluate before you buy
Start with the free sample, then a pilot
Request 500 images with annotations and full metadata from Nastaliq & Handwriting OCR — identical in format and quality to the paid corpus, no watermark, no sales call. Pilots are creditable against production and start within five business days of a signed SOW.
Average first response under 6 business hours. NDA signed same day.
- Consent form, licence text and DPA available before you commit
- Datasheet lists composition, QA method and known limitations
- Subsets by language, region or condition on request
- Custom collection to your spec using the same protocol
Answers
Questions about this dataset
Nastaliq has a sloping baseline, very long ligature context and glyphs that overlap vertically, so line and character segmentation approaches designed for Arabic fail. Our research post on Urdu OCR explains the mechanics.
Reading order across the page and field-level labels for structured documents such as forms and receipts.
A typical first delivery is 10k–100k pages; the free sample is a slice with the same annotation.