Python-PySpark Data Pipeline Support

Job ID: 40224779

Budget: ₹12,500 – ₹37,500 INR

I’m standing up a series of production data pipelines and need an IT professional who can move comfortably between Python scripting, SQL optimisation, PySpark transformations and Airflow orchestration. The immediate focus is end-to-end pipeline build-out: designing clean ingestion logic, transforming data in Spark, writing efficient queries and scheduling everything through well-structured Airflow DAGs.

If you can demonstrate hands-on experience across all four technologies - Python, SQL, PySpark and Airflow—and enjoy owning a pipeline from raw source to curated output, I’d like to work together.

Deliverables I’m expecting:
• A working set of PySpark jobs that handle ingestion, transformation and output staging
• Airflow DAGs that schedule, monitor and alert on each job’s success or failure
• Parameterised SQL/Python modules reusable across datasets
• Clear README-style documentation so another engineer can pick up the workflow without hand-holding

Code should be version-controlled (Git) and written with production reliability in mind: idempotent runs, informative logging and sensible error handling. Once the first pipelines are stable, there will be ongoing opportunities to tune performance, extend coverage and automate additional workflows.

Let me know your availability and a brief example of a pipeline you have recently delivered that shows these skills working together.