New: conversational Urdu and Punjabi speech corpora now licensable  See the catalog →

Image & video data · ready-made dataset

Nastaliq and handwriting OCR dataset

Nastaliq breaks OCR pipelines built for Arabic: sloping baselines, extreme ligature context and overlapping glyphs. Handwritten Nastaliq is harder again. This set pairs real printed and handwritten documents with reading-order and field-level annotation so a document-AI model can learn the script as it is actually written.

Free unwatermarked samplePer-record consentTraining-rights licence14-day acceptance

Datasheet summary

What is in Nastaliq & Handwriting OCR

Printed Nastaliq and handwritten forms, receipts, prescriptions and ledgers with reading-order and field-level annotation.

Dataset IDTDA-IV-002
ModalityImage
Languages / coverageUrdu, Sindhi, Shahmukhi Punjabi
Typical first delivery10k–100k pages
Free sample500 images with annotations and full metadata, within one business day
ProvenanceCollected in Pakistan; collection date, location, contributor pseudonym, device profile and consent version on every record
LicencePerpetual, worldwide, sublicensable; derivative-model rights stated explicitly. Exclusivity optional. Licensing detail
DeliveryYour S3, GCS or Azure bucket with manifests, checksums, QA report and datasheet
Guarantee14-day acceptance window; anything outside signed tolerance re-collected free

Included in every delivery

  • Printed Nastaliq and handwritten documents
  • Forms, receipts, prescriptions and ledgers
  • Reading-order annotation
  • Field-level annotation
  • Urdu, Sindhi and Shahmukhi Punjabi

Built for

  • Urdu, Sindhi and Shahmukhi Punjabi OCR
  • Handwriting recognition for prescriptions, forms and ledgers
  • Document-AI field extraction and layout understanding
  • Digitisation of paper records in South Asia

Evaluate before you buy

Start with the free sample, then a pilot

Request 500 images with annotations and full metadata from Nastaliq & Handwriting OCR — identical in format and quality to the paid corpus, no watermark, no sales call. Pilots are creditable against production and start within five business days of a signed SOW.

Average first response under 6 business hours. NDA signed same day.

  • Consent form, licence text and DPA available before you commit
  • Datasheet lists composition, QA method and known limitations
  • Subsets by language, region or condition on request
  • Custom collection to your spec using the same protocol

Answers

Questions about this dataset

Nastaliq has a sloping baseline, very long ligature context and glyphs that overlap vertically, so line and character segmentation approaches designed for Arabic fail. Our research post on Urdu OCR explains the mechanics.

Reading order across the page and field-level labels for structured documents such as forms and receipts.

A typical first delivery is 10k–100k pages; the free sample is a slice with the same annotation.

Free samples Get a quote