To do OCR (Optical Character Recognition) to extract text scanned in a PDF file
Budget: ₹12,500 – ₹37,500 INR
You have to extract people data (name, age, gender, etc) and a unique alphanumeric code from a PDF file. The PDF file is scanned & watermarked hence the OCR needs to be a little carefully handled. The OCR output can be sent to us in either text or excel format, so long as the EPIC is clearly extracted and is easily linkable to the name, we are good.
A sample screenshot from one of the pages of the PDF is attached.
I have encircled the unique alphanumeric codes.
The alphnumeric code has a standard template AAA9999999 which means 3 alpha and 7 numeric characters.
Use this template info to improve prediction accuracy.
Once you have decoded attached sample page, please count the unique alphanumeric codes correctly decoded, and mention the accuracy in your application to qualify for next step.
A sample screenshot from one of the pages of the PDF is attached.
I have encircled the unique alphanumeric codes.
The alphnumeric code has a standard template AAA9999999 which means 3 alpha and 7 numeric characters.
Use this template info to improve prediction accuracy.
Once you have decoded attached sample page, please count the unique alphanumeric codes correctly decoded, and mention the accuracy in your application to qualify for next step.