Automate Excel & PDF Extraction
Budget: $2 – $8 USD
I need a Python-based solution that pulls targeted information out of incoming Excel spreadsheets and PDF documents so I can stop copying values by hand. The script has to open each file, locate the fields I specify, and export the results in a clean, consolidated format ready for further analysis. Accuracy, repeatability, and clear logging are more important than fancy interfaces.
What I’m looking for
• A well-structured Python script (pandas, openpyxl, pdfplumber, or any comparable libraries you recommend) with helpful comments and a brief README.
• Robust handling of real-world quirks such as merged cells, variable column order, or slightly different PDF templates.
• Demonstrable testing so I can trust the extraction works the same way every time.
When you apply, please send a detailed project proposal that walks me through your approach, timeline, and any similar data-extraction work you’ve delivered. I’m specifically interested in the logic you plan to use and how you’ll validate edge cases. General résumés are less useful to me than a clear, step-by-step plan for this job.
What I’m looking for
• A well-structured Python script (pandas, openpyxl, pdfplumber, or any comparable libraries you recommend) with helpful comments and a brief README.
• Robust handling of real-world quirks such as merged cells, variable column order, or slightly different PDF templates.
• Demonstrable testing so I can trust the extraction works the same way every time.
When you apply, please send a detailed project proposal that walks me through your approach, timeline, and any similar data-extraction work you’ve delivered. I’m specifically interested in the logic you plan to use and how you’ll validate edge cases. General résumés are less useful to me than a clear, step-by-step plan for this job.
Related categories:
Python
Data Processing
Excel
Software Architecture
Data Extraction
Data Analysis
Automation
Pandas