Build a PDF OCR Pipeline using Tesseract

Job ID: 34480069

Budget: $30 – $250 USD

I'd like help building a PDF OCR application/back-end service (which will be used in an AWS EC2 environment). The objective of this application is to take input PDFs and perform OCR on them if they need it, returning a new PDF with the OCRed text as a layer.

The project can utilize existing tools like Tesseract or OCRPy (https://github.com/maxent-ai/ocrpy) or any other available open source tooling..