multimodal document classifier
Budget: $30 – $250 USD
Custom multimodal document classifier to accurately classify PDF documents using both text and image content. The documents can be single page or multi-page and are often combined in one package. For reference this approach looks good: https://medium.com/towards-artificial-intelligence/multimodal-deep-multipage-document-classification-using-both-image-and-text-629e5a2fdb47 (see the supporting article linked in that article). I have access to many labeled one-page documents with which I can train a model. Deliverable is a python notebook using all open source tools. For OCR I'm fine to use AWS textract or Google vision if easier and the performance is better, but would ideally use tesseract.