PySpark Glue Job for S3 to Redshift
Budget: $30 – $250 USD
I'm looking for a skilled PySpark developer to create a Glue Job that connects daily to an S3 bucket. The job should handle 5 different types of CSV files, converting them into Parquet format and subsequently ingesting them into a Redshift cluster using a provided DDL file for configuration.
Key Requirements:
- Use of best coding practices
- Implementing systems to avoid duplicates before data is ingested into Redshift
- Validating transfer with a final client repository
- Utilizing a specified client code library
- Converting CSV files to Parquet format
Ideal Skills and Experience:
- Proficiency in PySpark and AWS Glue
- Experience with S3 and Redshift
- Familiarity with CSV and Parquet file formats
- Strong understanding of data validation techniques
- Excellent coding skills with a focus on best practices
- Knowledge of data deduplication methods
Please note, the specific client code library to be used has not been disclosed yet, but will be provided to the selected freelancer. The job will also require the freelancer to check for duplicates and remove them prior to ingestion into Redshift.
Key Requirements:
- Use of best coding practices
- Implementing systems to avoid duplicates before data is ingested into Redshift
- Validating transfer with a final client repository
- Utilizing a specified client code library
- Converting CSV files to Parquet format
Ideal Skills and Experience:
- Proficiency in PySpark and AWS Glue
- Experience with S3 and Redshift
- Familiarity with CSV and Parquet file formats
- Strong understanding of data validation techniques
- Excellent coding skills with a focus on best practices
- Knowledge of data deduplication methods
Please note, the specific client code library to be used has not been disclosed yet, but will be provided to the selected freelancer. The job will also require the freelancer to check for duplicates and remove them prior to ingestion into Redshift.