PDF to DOCX converter using python which will support Bangla Bangladeshi language

Job ID: 36449469

Budget: $30 – $250 USD

I need a pdf2docx converter with source code that can accurately extract text, tables, and images from PDF files. It should preserve the original formatting, such as line alignment, tabs, and output the extracted information in a .doc format It will support Bangla and English language.

I will be happy if anyone helps me for making a pdf2docx converter.

Note:
I use pdf2docx which gives tables, line alignment, and tabs perfectly but problem is that the text is not correct because it does not support Bangla Bangladeshi Language.

To support Bangla Bangladeshi Language I use pytesseract OCR It gives text accurately but problem is that it does not support tables, line alignment, and tabs.

details link:
https://drive.google.com/file/d/1yDQ-hHaJQSmbUn608lutzo5bjLVKmSAx/view?usp=share_link
Related categories: Python Software Architecture OCR OpenCV