Build a PDF OCR Pipeline using Tesseract
Budget: $30 – $250 USD
I'd like help building a PDF OCR application/back-end service (which will be used in an AWS EC2 environment). The objective of this application is to take input PDFs and perform OCR on them if they need it, returning a new PDF with the OCRed text as a layer.
The project can utilize existing tools like Tesseract or OCRPy (https://github.com/maxent-ai/ocrpy) or any other available open source tooling..
The project can utilize existing tools like Tesseract or OCRPy (https://github.com/maxent-ai/ocrpy) or any other available open source tooling..