Optimize Spark Job on VERY Large Datasets (Data Enrichment)

Job ID: 35190623

Budget: €750 – €1,500 EUR

We have a Pipeline in which we are doing data enrichment. We aim to process a more than 30 TB data. Currently, we are only processing little less than 2percent (~200 GB) of that data, but we are facing memory issues.

Our Use case is very specific as we aim to do data enrichment on a huge, very huge dataset. Please read carefully attached document before sending your proposal.

Error 1
```
ExecutorLostFailure (executor 6 exited caused by one of the running tasks) Reason: Remote RPC client disassociated. Likely due to containers exceeding thresholds, or network issues. Check driver logs for WARN messages.
``

Error2
```
java.lang.OutOfMemoryError: Java heap space at org.apache.spark.sql.catalyst.expressions.UnsafeRow.copy(UnsafeRow.java:492)
``
Related categories: Linux Scala Spark Amazon S3 Apache Spark