Fix Azure Databricks Shuffle Errors and need training
Budget: ₹100 – ₹400 INR
I’m running data-transformation notebooks on Azure Databricks that read Parquet files and perform several heavy join operations. Under load the jobs crash with classic shuffle issues—“FetchFailedException”, “shuffle exceeded size”, and a few executor lost messages.
I need an experienced Databricks & Spark troubleshooter to:
• identify the root cause of these shuffle failures from the job and cluster logs,
• show (and document) the exact configuration, partitioning, or code changes that prevent the error,
• provide an optimised, tested notebook or script that completes the join workflow end-to-end without errors.
You will have remote access to the workspace (Python notebooks, cluster settings, Parquet sample data) and I’m happy to run any diagnostic commands you request. Success is a clean run of the current transformation pipeline plus a brief write-up of what was changed and why.
I need an experienced Databricks & Spark troubleshooter to:
• identify the root cause of these shuffle failures from the job and cluster logs,
• show (and document) the exact configuration, partitioning, or code changes that prevent the error,
• provide an optimised, tested notebook or script that completes the join workflow end-to-end without errors.
You will have remote access to the workspace (Python notebooks, cluster settings, Parquet sample data) and I’m happy to run any diagnostic commands you request. Success is a clean run of the current transformation pipeline plus a brief write-up of what was changed and why.