PySpark & SQL Hadoop Specialist
Budget: ₹1,500 – ₹12,500 INR
I have a Hadoop cluster holding several large data sets, and I need a seasoned PySpark developer who also writes rock-solid SQL. The immediate aim is to connect to the cluster (YARN/HDFS with Hive metastore), develop or refine PySpark jobs, optimise the accompanying SQL, and make sure everything runs smoothly end-to-end.
You’ll receive access to a staging namespace plus a sample of the data. Once the logic checks out we’ll promote the code to the full environment.
Deliverables
• A clean, well-commented PySpark notebook or .py job that executes successfully on the cluster
• The corresponding SQL script or view definitions ready for Hive or spark-sql
• A concise README detailing execution steps, parameters, and expected outputs
Acceptance criteria
• Jobs finish error-free on a 100 GB test slice
• Performance meets the runtime target we agree on before scaling up
• Output matches the sample I provide during onboarding
If you know Spark 3.x, HiveQL, and the nuances of tuning workloads on Hadoop, I look forward to working with you.
You’ll receive access to a staging namespace plus a sample of the data. Once the logic checks out we’ll promote the code to the full environment.
Deliverables
• A clean, well-commented PySpark notebook or .py job that executes successfully on the cluster
• The corresponding SQL script or view definitions ready for Hive or spark-sql
• A concise README detailing execution steps, parameters, and expected outputs
Acceptance criteria
• Jobs finish error-free on a 100 GB test slice
• Performance meets the runtime target we agree on before scaling up
• Output matches the sample I provide during onboarding
If you know Spark 3.x, HiveQL, and the nuances of tuning workloads on Hadoop, I look forward to working with you.