Python Script for Document Understanding

Job ID: 39291674

Budget: $10 – $30 CAD

I'm seeking a seasoned developer well-versed in Natural Language Processing (NLP) and Computer Vision. The project involves creating a clean, well-documented Python script utilizing Donut and LayoutLMv3 models for document comprehension.

This script's primary purpose is to automatically process scanned documents (like forms, invoices, certificates, and other types) from images, extract pertinent information, and categorize the document type based on its title.

Key Requirements:
- The script should be capable of handling various document types, including forms and invoices, as well as other documents.
- It should proficiently extract specific types of information such as text fields and tables from these documents.
- The extracted data should be outputted in JSON format.

Ideal Skills and Experience:
- Extensive experience in Python programming.
- In-depth knowledge of Natural Language Processing (NLP) techniques.
- Proficiency in Computer Vision applications.
- Familiarity with Donut and LayoutLMv3 models.
- Previous experience in creating well-documented scripts.
- Ability to handle and extract information from diverse document types.
Related categories: Python Software Architecture JSON Computer Vision NLP