PDF Image Segmentation & Text Extraction

Job ID: 38318274

Budget: $30 – $250 USD

I'm in need of a Python expert to create a custom solution for extracting text + math formulas and tables from PDF files. The data should then be assembled into a readable format (latex) to be stored in an SQLite database. The project should utilize software from existing open source projs.

Key Requirements:
- Coding a Python using open source components for OCR, pypdf, and Nougat to extract text and images seamlessly from PDF files.
- brief description of project details is attached

Ideal Skills and Experience:
- Proficiency in Python programming, specifically working with OCR and PDF libraries.
- Prior experience with image segmentation and text extraction from PDF files is highly desired.
- Usage of Github Copilot or equivalent tools is a plus.

The SQLite DB model/file will be provided
Related categories: Python Data Processing PDF OCR