Scanned eBook Extraction to Excel
Budget: ₹1,250 – ₹2,500 INR
I’m sitting on a fully scanned eBook that mixes paragraphs, headings, images and their captions on every page. What I need is a clean, structured Excel workbook that holds absolutely everything—both the text and the images—pulled straight from those scans.
Here’s the flow I’m envisioning: you run the pages through reliable OCR (ABBYY FineReader, Tesseract or whatever you trust for high accuracy), capture every block of text without losing line breaks or special characters, and export each embedded image at its original resolution. In the Excel file, place the text in clearly labelled columns (page number, heading, body text, caption) and give each extracted image an easy-to-follow file name that’s referenced in its own column. I’ll need the actual image files delivered alongside the spreadsheet, neatly organised in folders that mirror the page numbers.
Deliverables
• One Excel workbook containing all extracted text and image references
• A folder structure of the corresponding image files in original quality
I’ll provide the scanned PDF as soon as we start, along with a short sample so we can agree on layout. Accuracy is everything: spelling should match the source, images must not be down-scaled, and the page order has to stay intact. Let me know which OCR tool you plan to use and your estimated turnaround time.
Here’s the flow I’m envisioning: you run the pages through reliable OCR (ABBYY FineReader, Tesseract or whatever you trust for high accuracy), capture every block of text without losing line breaks or special characters, and export each embedded image at its original resolution. In the Excel file, place the text in clearly labelled columns (page number, heading, body text, caption) and give each extracted image an easy-to-follow file name that’s referenced in its own column. I’ll need the actual image files delivered alongside the spreadsheet, neatly organised in folders that mirror the page numbers.
Deliverables
• One Excel workbook containing all extracted text and image references
• A folder structure of the corresponding image files in original quality
I’ll provide the scanned PDF as soon as we start, along with a short sample so we can agree on layout. Accuracy is everything: spelling should match the source, images must not be down-scaled, and the page order has to stay intact. Let me know which OCR tool you plan to use and your estimated turnaround time.
Related categories:
Data Processing
Data Entry
Excel
PDF
OCR
Image Processing
Data Extraction
ABBYY FineReader