Optical Character Recognition from a scanned PDF file with people data

Job ID: 35110805

Budget: ₹12,500 – ₹37,500 INR

We have scanned PDF files that have thousands of people data with printed name, age, gender with a unique alphanumeric code. All these are positioned inside pre formatted rectangular ticket size boxes and these boxes are evenly spaced across all pages.

You have to do OCR on the PDF files and hand us the output in either of the following format.
1. Plain continuous text (easier option)
2. Excel format after cleaning up and putting the names, age, gender, unique code in separate columns. (tougher option).

We will do a test run on one PDF file and if that has a good accuracy - atleast 90%, we will go ahead with the rest of the files.

Screenshot of a sample page of the PDF document is uploaded for inspection.
This is a png file but actual files will be PDF.

Please report success on your OCR on attached picture, with percentage correct EPIC numbers while applying to the post.
Related categories: Python R Programming Language Data Science