Databricks Ingestion Bug Fix

Job ID: 40538379

Budget: £10 – £20 GBP

Our production workspace is throwing errors the moment new SQL-based tables hit the ingestion layer. Downstream jobs remain healthy, so the issue is isolated to the first step of the pipeline that pulls structured data into Databricks before any transformations begin.

I need you to jump into the existing notebooks and jobs, trace the root cause, and deliver a clean, repeatable fix. The environment runs on Databricks Runtime 12.x with Auto Loader feeding into Delta tables, so familiarity with Spark, PySpark, SQL, and cluster-level configs is essential. I will grant you workspace access and point you to the failing job runs and relevant logs.

Acceptance criteria:
• Ingestion job completes successfully three runs in a row with identical input files.
• No data loss or duplication in the target Delta tables (row counts match source).
• A concise summary of what caused the failure and how you corrected it, committed to our repo’s README.

Once the bug is resolved we can discuss optimising the rest of the pipeline, but first priority is getting ingestion back to green.