OCR Expert for PDF Data Extraction Improvement

Job ID: 40529921

Budget: $30 – $250 USD

Senior OCR / PDF Data Extraction Expert Needed

We are looking for an experienced developer with strong expertise in:

* OCR systems
* PDF parsing
* Table extraction
* Document AI
* Data normalization
* Matching and reconciliation algorithms

Project Overview

We have an existing web application that processes business documents (PDF invoices and related documents) and extracts structured line-item data.

The current system is already operational and includes:

* OCR pipeline
* Data extraction
* Matching engine
* Web interface

However, we are experiencing accuracy issues in specific document layouts, including:

* Incorrect table reconstruction
* Merged rows or merged numeric values
* Quantity/price parsing errors
* Duplicate line detection issues
* Aggregation inconsistencies
* Matching inaccuracies between related documents

What We Need

An expert who can:

1. Review the current extraction workflow.
2. Analyze problematic PDF samples.
3. Identify the root cause of calculation discrepancies.
4. Recommend and/or implement improvements.
5. Advise whether advanced AI-based table reconstruction is actually required or whether the issues can be solved through improved parsing logic.

Requirements

* Proven experience with OCR technologies.
* Experience with Google Vision, Azure Document Intelligence, AWS Textract, Tesseract, or similar.
* Strong knowledge of PDF table extraction.
* Experience debugging large-scale document-processing systems.
* Ability to review existing code and provide architectural recommendations.

Preferred

* Experience with invoice processing systems.
* Experience with financial/business document workflows.
* Experience with LLM-assisted document reconstruction.

Please include:

* Relevant projects.
* OCR/document AI experience.
* Suggested approach for diagnosing extraction accuracy issues.