Custom Document-to-Excel Conversion System
Budget: $15 – $25 USD
Job Post: Build Custom Document-to-Excel Conversion Pipeline
Project Overview
I’m looking for a developer or small team to build a pipeline that takes in multiple file formats (Excel, PowerPoint, PDF, images) and outputs the extracted information into a standardized Excel template/schema that I define.
This will be a custom hybrid solution (not just an off-the-shelf converter). The goal is to automate and standardize a manual process where I currently copy/paste or reformat files into Excel.
---
Key Requirements
Strong Python development skills.
Experience with document parsing librarie (e.g., `pdfplumber, Camelot, python-pptx, pandas openpyxl
OCR skills for scanned PDFs and images (Tesseract, Google Vision, AWS Textract, or Azure Form Recognizer).
Ability to design schema mapping logic (input →normalized format → target Excel schema).
Experience building ETL pipelines or automation scripts that process batches of files.
Nice to have Ability to wrap the workflow in a simple web app or folder-watcher automation so it’s user-friendly.
---
**Deliverables**
1. Scripts/pipeline that:
Ingests multiple file types (Excel, PowerPoint, PDFs, images).
*Extracts structured data (tables/text).
Maps it into a defined Excel schema.
Outputs a clean, validated Excel file.
2. Documentation so I can adjust schema mapping later.
3. Bonus: optional lightweight UI (upload → download result) or automation
Project Overview
I’m looking for a developer or small team to build a pipeline that takes in multiple file formats (Excel, PowerPoint, PDF, images) and outputs the extracted information into a standardized Excel template/schema that I define.
This will be a custom hybrid solution (not just an off-the-shelf converter). The goal is to automate and standardize a manual process where I currently copy/paste or reformat files into Excel.
---
Key Requirements
Strong Python development skills.
Experience with document parsing librarie (e.g., `pdfplumber, Camelot, python-pptx, pandas openpyxl
OCR skills for scanned PDFs and images (Tesseract, Google Vision, AWS Textract, or Azure Form Recognizer).
Ability to design schema mapping logic (input →normalized format → target Excel schema).
Experience building ETL pipelines or automation scripts that process batches of files.
Nice to have Ability to wrap the workflow in a simple web app or folder-watcher automation so it’s user-friendly.
---
**Deliverables**
1. Scripts/pipeline that:
Ingests multiple file types (Excel, PowerPoint, PDFs, images).
*Extracts structured data (tables/text).
Maps it into a defined Excel schema.
Outputs a clean, validated Excel file.
2. Documentation so I can adjust schema mapping later.
3. Bonus: optional lightweight UI (upload → download result) or automation