PDF Ledger-to-Excel Conversion
Budget: ₹12,500 – ₹37,500 INR
I have a lengthy financial ledger locked inside a poorly formatted PDF. Irregular spacing, line breaks in the middle of numbers, and random text strings currently prevent any meaningful accounting work or analysis.
Your task is to pull every record out of that PDF, standardise the layout, and deliver a single, well-structured Excel workbook that is ready for pivot-tables, reconciliations, or import into bookkeeping software.
Key points to keep in mind
• Source: one multi-page PDF containing chronological ledger entries
• Common issues: merged columns, wrapped descriptions, inconsistent date and currency formats, stray headers/footers
• End format: native .xlsx with clean column headers and one row per transaction
Deliverables (all mandatory)
1. An Excel file with fully normalised data—no empty spacer rows, no merged cells, uniform date and amount formatting.
2. A concise processing log (plain text or a second sheet) summarising any assumptions, transformations, or fields that could not be recovered verbatim.
I will spot-check totals against the original PDF, so accuracy is critical. Automated extraction tools are fine as long as the final result passes visual inspection and basic sum checks. If you rely on Python, Power Query, or similar, feel free to include your script; that’s optional but appreciated.
Your task is to pull every record out of that PDF, standardise the layout, and deliver a single, well-structured Excel workbook that is ready for pivot-tables, reconciliations, or import into bookkeeping software.
Key points to keep in mind
• Source: one multi-page PDF containing chronological ledger entries
• Common issues: merged columns, wrapped descriptions, inconsistent date and currency formats, stray headers/footers
• End format: native .xlsx with clean column headers and one row per transaction
Deliverables (all mandatory)
1. An Excel file with fully normalised data—no empty spacer rows, no merged cells, uniform date and amount formatting.
2. A concise processing log (plain text or a second sheet) summarising any assumptions, transformations, or fields that could not be recovered verbatim.
I will spot-check totals against the original PDF, so accuracy is critical. Automated extraction tools are fine as long as the final result passes visual inspection and basic sum checks. If you rely on Python, Power Query, or similar, feel free to include your script; that’s optional but appreciated.
Related categories:
Visual Basic
Data Processing
Data Entry
Excel
Data Cleansing
Data Extraction
Data Analysis
Data Management