Voter List (Electoral Roll) PDF Data Extraction – Scanned + Multi-Language
Budget: ₹600 – ₹1,500 INR
I need a freelancer to make a script which can extract voter list (electoral roll) PDF data into a structured Excel/CSV format.
These PDFs can be in multiple languages (Hindi/English/Malayalam/Telugu/Marathi etc.) and may have different layouts/styles (grid cards, 3x10 card pages, scanned PDFs, mixed quality).
I want a solution that works per PDF style and can process multiple PDFs from a folder.
Scope of Work
✅ Convert each voter “card” into a row in Excel/CSV
✅ Handle multi-language text properly
✅ Work with scanned/image PDFs (OCR may be required)
✅ Maintain correct columns and consistent formatting
✅ Option to generate one combined file for all PDFs + separate per PDF output
Required Fields (Typical)
Serial Number
EPIC / Voter ID
Name
Relation type + relation name (Father/Mother/Husband etc.)
House No.
Age
Gender
Section/Part info and top header details (if present: AC name/number, section name/number)
Output
Excel (.xlsx) and/or CSV
Clean columns + proper headers
Missing values should be blank (don’t guess)
Must Have Skills
Python / OCR / PDF parsing experience
Strong with layout-based extraction (card/grid detection)
Experience with Indian voter list/electoral roll PDFs is a big plus
Ability to show sample output from my provided sample PDF
Proof of Capability (Important)
Before awarding full work, I will share 1–2 sample PDFs.
You must provide 10 sample extracted rows to confirm accuracy.
Deliverables
1. Final extracted Excel/CSV (sample)
2. Script/tool (preferred) that can run on Windows and process folder input → output folder
3. Short usage instructions
These PDFs can be in multiple languages (Hindi/English/Malayalam/Telugu/Marathi etc.) and may have different layouts/styles (grid cards, 3x10 card pages, scanned PDFs, mixed quality).
I want a solution that works per PDF style and can process multiple PDFs from a folder.
Scope of Work
✅ Convert each voter “card” into a row in Excel/CSV
✅ Handle multi-language text properly
✅ Work with scanned/image PDFs (OCR may be required)
✅ Maintain correct columns and consistent formatting
✅ Option to generate one combined file for all PDFs + separate per PDF output
Required Fields (Typical)
Serial Number
EPIC / Voter ID
Name
Relation type + relation name (Father/Mother/Husband etc.)
House No.
Age
Gender
Section/Part info and top header details (if present: AC name/number, section name/number)
Output
Excel (.xlsx) and/or CSV
Clean columns + proper headers
Missing values should be blank (don’t guess)
Must Have Skills
Python / OCR / PDF parsing experience
Strong with layout-based extraction (card/grid detection)
Experience with Indian voter list/electoral roll PDFs is a big plus
Ability to show sample output from my provided sample PDF
Proof of Capability (Important)
Before awarding full work, I will share 1–2 sample PDFs.
You must provide 10 sample extracted rows to confirm accuracy.
Deliverables
1. Final extracted Excel/CSV (sample)
2. Script/tool (preferred) that can run on Windows and process folder input → output folder
3. Short usage instructions
Related categories:
Business, Accounting, Human Resources & Legal
Python
OCR
Data Scraping
Data Extraction