multimodal document classifier

Job ID: 36963972

Budget: $30 – $250 USD

Custom multimodal document classifier to accurately classify PDF documents using both text and image content. The documents can be single page or multi-page and are often combined in one package. For reference this approach looks good: https://medium.com/towards-artificial-intelligence/multimodal-deep-multipage-document-classification-using-both-image-and-text-629e5a2fdb47 (see the supporting article linked in that article). I have access to many labeled one-page documents with which I can train a model. Deliverable is a python notebook using all open source tools. For OCR I'm fine to use AWS textract or Google vision if easier and the performance is better, but would ideally use tesseract.
Related categories: Python Machine Learning (ML)