Processing 15 GB text file using GCP Dataproc with Spark (fixed width files)
Budget: ₹1,500 – ₹12,500 INR
I have a pyspark code which is working for small files to process fixed width files on GCP dataproc cluster, but when I'm reading 15GB of compressed gzip text file, it is taking time to either save/load in BigQuery table and unable to fix this issue. Need someone to identify the root cause of this and resolved this issue