OCR Expert for PDF Data Extraction Improvement
Budget: $30 – $250 USD
Senior OCR / PDF Data Extraction Expert Needed
We are looking for an experienced developer with strong expertise in:
* OCR systems
* PDF parsing
* Table extraction
* Document AI
* Data normalization
* Matching and reconciliation algorithms
Project Overview
We have an existing web application that processes business documents (PDF invoices and related documents) and extracts structured line-item data.
The current system is already operational and includes:
* OCR pipeline
* Data extraction
* Matching engine
* Web interface
However, we are experiencing accuracy issues in specific document layouts, including:
* Incorrect table reconstruction
* Merged rows or merged numeric values
* Quantity/price parsing errors
* Duplicate line detection issues
* Aggregation inconsistencies
* Matching inaccuracies between related documents
What We Need
An expert who can:
1. Review the current extraction workflow.
2. Analyze problematic PDF samples.
3. Identify the root cause of calculation discrepancies.
4. Recommend and/or implement improvements.
5. Advise whether advanced AI-based table reconstruction is actually required or whether the issues can be solved through improved parsing logic.
Requirements
* Proven experience with OCR technologies.
* Experience with Google Vision, Azure Document Intelligence, AWS Textract, Tesseract, or similar.
* Strong knowledge of PDF table extraction.
* Experience debugging large-scale document-processing systems.
* Ability to review existing code and provide architectural recommendations.
Preferred
* Experience with invoice processing systems.
* Experience with financial/business document workflows.
* Experience with LLM-assisted document reconstruction.
Please include:
* Relevant projects.
* OCR/document AI experience.
* Suggested approach for diagnosing extraction accuracy issues.
We are looking for an experienced developer with strong expertise in:
* OCR systems
* PDF parsing
* Table extraction
* Document AI
* Data normalization
* Matching and reconciliation algorithms
Project Overview
We have an existing web application that processes business documents (PDF invoices and related documents) and extracts structured line-item data.
The current system is already operational and includes:
* OCR pipeline
* Data extraction
* Matching engine
* Web interface
However, we are experiencing accuracy issues in specific document layouts, including:
* Incorrect table reconstruction
* Merged rows or merged numeric values
* Quantity/price parsing errors
* Duplicate line detection issues
* Aggregation inconsistencies
* Matching inaccuracies between related documents
What We Need
An expert who can:
1. Review the current extraction workflow.
2. Analyze problematic PDF samples.
3. Identify the root cause of calculation discrepancies.
4. Recommend and/or implement improvements.
5. Advise whether advanced AI-based table reconstruction is actually required or whether the issues can be solved through improved parsing logic.
Requirements
* Proven experience with OCR technologies.
* Experience with Google Vision, Azure Document Intelligence, AWS Textract, Tesseract, or similar.
* Strong knowledge of PDF table extraction.
* Experience debugging large-scale document-processing systems.
* Ability to review existing code and provide architectural recommendations.
Preferred
* Experience with invoice processing systems.
* Experience with financial/business document workflows.
* Experience with LLM-assisted document reconstruction.
Please include:
* Relevant projects.
* OCR/document AI experience.
* Suggested approach for diagnosing extraction accuracy issues.
Related categories:
Data Processing
Algorithm
OCR
Data Extraction
Data Analysis
Data Integration
Data Management
AI Development