Large-Scale PDF Data Extraction Specialist

Job ID: 39659726

Budget: ₹1,500 – ₹12,500 INR

We are seeking a freelancer to assist with a large-scale data extraction project. You will work with a set of 5,000 PDF documents, extracting tables, graphs, and performing curve fitting, followed by validation of the results.

Key Responsibilities:

Run OCR on PDF files to extract textual and graphical data.

Extract tables and graphs from each PDF using provided scripts (modification or minor adaptation may occasionally be required to handle variations between different PDF formats).

Perform curve fitting on data extracted from graphs as specified.

Validate the extracted data for accuracy and consistency.

Communicate any edge cases or difficulties with problematic PDFs for collaborative troubleshooting.

Keep meticulous records of processed and validated files.

Requirements:

Proven experience with leading OCR tools (e.g., Tesseract, Adobe, ABBYY, etc.).

Familiarity with automated table and graph extraction methods from PDFs.

Experience with curve fitting using Python (e.g., NumPy, SciPy) or similar libraries.

Ability to understand and slightly tweak existing code/scripts to improve extraction for varied PDFs.

High attention to detail and dedication to thorough validation.

Capable of handling repetitive tasks methodically for large datasets.

Scope:

5,000 unique PDFs to be extracted, fitted, and validated.

Tools Provided: Existing scripts for OCR, table/graph extraction, and curve fitting. Candidate may be asked to adjust these scripts as needed for edge cases.

If you have a keen eye for detail, experience with data extraction, and are ready for a high-volume project, we’d love to hear from you! Please share examples of similar projects you’ve completed.

Note: The requirement focuses on ability to use tools and scripts provided for OCR and extraction rather than deep OCR expertise.