OCR text extraction from 700 JPG pages
Budget: $200 – $500 AUD
I have approximately 700 pages, in JPG format, containing horse pedigrees that I need to extract into CSV files. An example is attached. The information to be extracted includes:
- Detect whether the header is blue or red
- Extract the fields in the header (name, date of birth, breeder, etc)
- Extract the first four columns in the main body (the fifth column is not required)
- Extract the table of progeny at the bottom, including the sex which needs to be parsed from the icon (either male, female or scissor symbol)
In the attachment you'll see an annotated example and a list of fields.
The successful freelancer will be provided with the 700 pages as JPGs and will be required to extract all relevant information into a CSV file.
- Detect whether the header is blue or red
- Extract the fields in the header (name, date of birth, breeder, etc)
- Extract the first four columns in the main body (the fifth column is not required)
- Extract the table of progeny at the bottom, including the sex which needs to be parsed from the icon (either male, female or scissor symbol)
In the attachment you'll see an annotated example and a list of fields.
The successful freelancer will be provided with the 700 pages as JPGs and will be required to extract all relevant information into a CSV file.