PDF Table OCR to CSV
Budget: $30 – $250 USD
I have a scanned-image PDF that contains several pages of simple tables. Although the structure stays mostly the same, a few pages show small layout variations, so the solution has to cope with that gracefully. Every column that appears on the page must be captured—no selective extraction—then written to a clean, comma-separated CSV that preserves row order.
You are free to use any reliable OCR and tabular-extraction stack (Tesseract, AWS Textract, Tabula, Camelot, or a custom Python script are all fine so long as the final data are accurate). What matters is that the numbers and text lines up correctly in the resulting file.
Deliverables:
• One CSV file containing all table data, ready to open in Excel or import into a database.
• A brief note on the toolchain or code used so I can reproduce the process if the PDF is updated later.
Accuracy is more important than speed, so feel free to build in verification steps to double-check unusual rows caused by the minor layout changes. Let me know your approach and the timeframe you’ll need.
You are free to use any reliable OCR and tabular-extraction stack (Tesseract, AWS Textract, Tabula, Camelot, or a custom Python script are all fine so long as the final data are accurate). What matters is that the numbers and text lines up correctly in the resulting file.
Deliverables:
• One CSV file containing all table data, ready to open in Excel or import into a database.
• A brief note on the toolchain or code used so I can reproduce the process if the PDF is updated later.
Accuracy is more important than speed, so feel free to build in verification steps to double-check unusual rows caused by the minor layout changes. Let me know your approach and the timeframe you’ll need.