Data Analyst or Python Developer Needed
Budget: $250 – $750 USD
TIME SENSITIVE PROJECT
Description:
I have approximately 200 reports in Excel and PDF format. These reports contain tables or structured/semi-structured data, but the formatting, field names, and file naming conventions vary significantly across files.
I'm looking for a skilled data analyst or Python developer who can help me compare these reports and identify which ones are at least 60% similar in content. This will require fuzzy matching techniques and possibly data normalization.
Responsibilities:
Extract data from PDF and Excel reports (some may require OCR or table parsing).
Clean and normalize the data across all files.
Compare the reports and determine which are ≥60% similar based on data content.
Deliver a summary of matched report pairs or groups with similarity scores.
(Optional) Provide a Python script or lightweight tool so I can run this again in the future.
Skills Needed:
Python (pandas, fuzzywuzzy, difflib, or NLP libraries)
Excel data parsing (openpyxl, xlrd)
PDF parsing (pdfplumber, PyPDF2, or Camelot)
Experience with fuzzy matching and data cleaning
Ability to handle unstructured, messy datasets
Deliverables:
A report or spreadsheet listing grouped/matched reports by similarity level
A similarity matrix (optional)
A Python script or notebook to reproduce or re-run the process (optional but preferred)
To Apply:Please briefly describe:
Relevant experience or similar projects you've done
Estimated time/cost to complete the project
Description:
I have approximately 200 reports in Excel and PDF format. These reports contain tables or structured/semi-structured data, but the formatting, field names, and file naming conventions vary significantly across files.
I'm looking for a skilled data analyst or Python developer who can help me compare these reports and identify which ones are at least 60% similar in content. This will require fuzzy matching techniques and possibly data normalization.
Responsibilities:
Extract data from PDF and Excel reports (some may require OCR or table parsing).
Clean and normalize the data across all files.
Compare the reports and determine which are ≥60% similar based on data content.
Deliver a summary of matched report pairs or groups with similarity scores.
(Optional) Provide a Python script or lightweight tool so I can run this again in the future.
Skills Needed:
Python (pandas, fuzzywuzzy, difflib, or NLP libraries)
Excel data parsing (openpyxl, xlrd)
PDF parsing (pdfplumber, PyPDF2, or Camelot)
Experience with fuzzy matching and data cleaning
Ability to handle unstructured, messy datasets
Deliverables:
A report or spreadsheet listing grouped/matched reports by similarity level
A similarity matrix (optional)
A Python script or notebook to reproduce or re-run the process (optional but preferred)
To Apply:Please briefly describe:
Relevant experience or similar projects you've done
Estimated time/cost to complete the project