Need Python developer to Develop a script to OCR and ALTO to generate pdf and other output

Job ID: 32723442

Budget: $30 – $250 USD

The main scope of the project is to create a python bash or perl script that process OCR over a master folder of TIFF files or a PDF file.

The workflow will be:

1. Read a Master directory file with .Tiff files uncompressed

2. Start OCR processing

3. Generate Output files:

3.1. Output #1: Master PDF 1.4 (PDF/A-1)
3.2. Output #2: Access PDF 1.4 (PDF/A-1) - (Never bigger than 20mb.)
3.3. Output #2: ALTO XML with all the OCR information
3.4. Output #3: TXT with all the OCR text results
3.5. Output #4: TXT with all the OCR text results

4. Create XML with the info of the process
4.1 XML has to record all the processes and also record any failure on the process.

The project scope is to generate a script on bash, perl or python based on any Open-source OCR tools to do those tasks. To complete the project, the script has to be tested, checked and documented in detail. No chance to make anything different. During the development, the developer has to give us a report of the fields and values we have to work with and we will help on the definition. Can be a long term project if the results are good with much more image processing tasks.

Some valid Open-source tools that can bu used are:
- http://kraken.re/master/index.html
- https://pypi.org/project/pytesseract/
- https://github.com/tesseract-ocr/tesseract

Reference documentation links:
- https://github.com/altoxml
- https://www.loc.gov/standards/alto/
- https://www.loc.gov/standards/alto/techcenter/structure.html

We can provide image samples and also give much more detail of the script to develop.

Other options can be used but need to be validated by us before start the development.
Related categories: Perl XML Python Software Architecture Shell Script