Comprehensive Mixed Data Cleansing

Job ID: 40452291

Budget: ₹100 – ₹400 INR

I have a single dataset that blends both free-form text and numerical fields, and it needs to be made analysis-ready.

For the text columns I need every duplicate row removed, all formatting inconsistencies ironed out, and obvious spelling mistakes corrected so the wording is uniform and machine-readable.

On the numerical side I want clear, documented treatment of outliers, sensible imputation of missing values, and a consistent unit scale across comparable measures.

You are free to work in Python (Pandas, NumPy, open-source spell-check libraries), R, or another reliable toolset as long as the results can be reproduced. Please include a brief outline of the cleaning logic in well-commented code or a notebook so I can audit each decision.

Deliverables
• Cleaned master dataset in its original file format
• Reproducible script or notebook with explanatory comments
• One-page summary detailing what was changed and why

I will provide the raw file and any relevant data dictionaries once we begin; just let me know the preferred format.