Python Developer for Optimized Data Transformation on Large Datasets (Time Series & Position Data)

Job ID: 38616984

Budget: ₹1,500 – ₹12,500 INR

We are seeking an experienced Python developer skilled in handling large time-series and positional data for a high-performance data processing pipeline. The task requires optimized and parallelized processing of reading massive CSV, gzipped CSV or Parquet files and transforming/processing the data. The developer must ensure efficient handling, grouping, sorting, and imputation of data, as well as implementation of advanced data bucketing strategies. The project also requires robust error-handling mechanisms, including the ability to track progress and resume operations after a crash or interruption without duplicating previously processed data.

Requirements:

Expertise in Python, especially libraries like Pandas, Dask, or PySpark for parallel processing.
Experience with time-series data processing and geospatial data.
Proficiency in working with large datasets (several gigabytes to terabytes).
Knowledge of efficient I/O operations with CSV/Parquet formats.
Experience with error recovery and progress tracking in data pipelines.
Ability to write clean, optimized, and scalable code.

Please provide examples of similar projects you have worked on and how you ensured performance optimization.
Related categories: Python Data Processing Big Data