Text Extraction from PDF & Image Files
Budget: $750 – $1,500 USD
I need a professional freelancer to extract text from both PDF and Image files. The extracted text should be formatted into a CSV or Excel file. The documents contain printed text, tables, and charts.
Ideal Skills:
- Expertise in OCR, Deep Learning and Vision Model data extraction techniques
- Proficient in using open source software for PDF and image text extraction
Requirements:
- Build an utility for taking a document as input and extract attributes as predefined as output. The exctracted fields should have the accuracy of extarction.
- All open source packages/softwares/LLMs
- The utility should be able to handle documents with different layouts
- The utility should have learning capability like models can be retrained with new corpus
- The processing time for a single page document should be within few seconds (5 Sec)
- The utility should be able to process lower quality image/pdf with fair extraction accuracy
Ideal Skills:
- Expertise in OCR, Deep Learning and Vision Model data extraction techniques
- Proficient in using open source software for PDF and image text extraction
Requirements:
- Build an utility for taking a document as input and extract attributes as predefined as output. The exctracted fields should have the accuracy of extarction.
- All open source packages/softwares/LLMs
- The utility should be able to handle documents with different layouts
- The utility should have learning capability like models can be retrained with new corpus
- The processing time for a single page document should be within few seconds (5 Sec)
- The utility should be able to process lower quality image/pdf with fair extraction accuracy
Related categories:
OCR
Image Processing
Deep Learning
AI Image-to-text
Large Language Models (LLMs)