Pdf to Doc Converter Using OCR
Budget: $30 – $250 USD
I am looking for a Python programmer to help me create a PDF to DOCX converter using OCR technology.
The software should be able to accurately extract text, tables, fonts, font sizes, bold and italic formatting, as well as images from PDF files.
Furthermore, it should preserve the original formatting, such as line alignment, tabs, and other elements, and output the extracted information in a .doc format with identical formatting to the original PDF file.
It will support Bangla and English language.
The software should be able to accurately extract text, tables, fonts, font sizes, bold and italic formatting, as well as images from PDF files.
Furthermore, it should preserve the original formatting, such as line alignment, tabs, and other elements, and output the extracted information in a .doc format with identical formatting to the original PDF file.
It will support Bangla and English language.