Python utility to create a PDF file from image files, with OCR layer

Job ID: 37332552

Budget: $750 – $1,500 USD

Develop a windows command line Python application, that creates one PDF file from all existing images withing a directory and adds an invisible text layer (OCR extracted text) for each page.

Images files will be in local disk.

OCR extracted text will be obtained from “Microsoft Azure / Document Intelligence / FormRecognizer / Layout” API. Returned “.json” file has a list of all text fragments e box positions. These text fragments should be put near corresponding words positions in images.

In attached files, there is a detailed description of the task (.docx file), including test files.
Related categories: Python Azure