Text Extraction from PDF Files - 11/09/2024 18:42 EDT

Job ID: 38563058

Budget: $250 – $750 AUD

Need to extract text from PDF files which include images and save the data in an Excel spreadsheet.

Some of the data to be extracted appears as text on a white background. Some data is text superimposed onto an image. Each PDF file has three images, the text superimposed on each image is identical.

I imagine this as an OCR data processing exercise rather than transcription so that once a template has been developed it can be applied easily to mutiple files. The solution can be in the form of an app that I can use on my desktop to process files in small batches. There are a large number of PDF files so it is impractical to upload them for processing by a third party.

This is a once only project so the app is not required to be suitable for sharing with other users. It will be used on a Mac OS.

Ideal skills for this task include:
- Experience with OCR (Optical Character Recognition) software
- Experience with Google Cloud Vision OCR
- Attention to detail to ensure accuracy in text extraction
- Familiarity with Excel for data presentation

I have attached the following to illustrate the project:
1. A marked up PDF file showing the required data fields
2. An Excel file sample to be used for data presentation
3. A representative PDF file
Related categories: Data Processing Data Entry Excel Web Scraping OCR