Fix Azure Databricks Shuffle Errors and need training

Job ID: 40166493

Budget: ₹100 – ₹400 INR

I’m running data-transformation notebooks on Azure Databricks that read Parquet files and perform several heavy join operations. Under load the jobs crash with classic shuffle issues—“FetchFailedException”, “shuffle exceeded size”, and a few executor lost messages.

I need an experienced Databricks & Spark troubleshooter to:
• identify the root cause of these shuffle failures from the job and cluster logs,
• show (and document) the exact configuration, partitioning, or code changes that prevent the error,
• provide an optimised, tested notebook or script that completes the join workflow end-to-end without errors.

You will have remote access to the workspace (Python notebooks, cluster settings, Parquet sample data) and I’m happy to run any diagnostic commands you request. Success is a clean run of the current transformation pipeline plus a brief write-up of what was changed and why.
Related categories: Python Data Processing Azure Spark ETL Microsoft Azure