Text Extraction from PDF & Image Files

Job ID: 39088173

Budget: $750 – $1,500 USD

I need a professional freelancer to extract text from both PDF and Image files. The extracted text should be formatted into a CSV or Excel file. The documents contain printed text, tables, and charts.

Ideal Skills:
- Expertise in OCR, Deep Learning and Vision Model data extraction techniques
- Proficient in using open source software for PDF and image text extraction

Requirements:
- Build an utility for taking a document as input and extract attributes as predefined as output. The exctracted fields should have the accuracy of extarction.
- All open source packages/softwares/LLMs
- The utility should be able to handle documents with different layouts
- The utility should have learning capability like models can be retrained with new corpus
- The processing time for a single page document should be within few seconds (5 Sec)
- The utility should be able to process lower quality image/pdf with fair extraction accuracy