PySpark & SQL Hadoop Specialist

Job ID: 40227186

Budget: ₹1,500 – ₹12,500 INR

I have a Hadoop cluster holding several large data sets, and I need a seasoned PySpark developer who also writes rock-solid SQL. The immediate aim is to connect to the cluster (YARN/HDFS with Hive metastore), develop or refine PySpark jobs, optimise the accompanying SQL, and make sure everything runs smoothly end-to-end.

You’ll receive access to a staging namespace plus a sample of the data. Once the logic checks out we’ll promote the code to the full environment.

Deliverables
• A clean, well-commented PySpark notebook or .py job that executes successfully on the cluster
• The corresponding SQL script or view definitions ready for Hive or spark-sql
• A concise README detailing execution steps, parameters, and expected outputs

Acceptance criteria
• Jobs finish error-free on a 100 GB test slice
• Performance meets the runtime target we agree on before scaling up
• Output matches the sample I provide during onboarding

If you know Spark 3.x, HiveQL, and the nuances of tuning workloads on Hadoop, I look forward to working with you.