Data extraction from pdf files using OCR + python

Job ID: 34480253

Budget: $15 – $25 USD

The job consists in extracting data from a series of yearly publications reporting information on the organization of US local newspapers.

The initial job is for four volumes of about 100 pages each. However, it will eventually extend to an additional 6-7 volumes of similar length.

Attached is a sample page (in pdf) to get a sense of the task. The format is very much stable within and across publications.
Related categories: C Programming Python Data Processing OCR