Data Engineering Specialist Needed

Job ID: 38088526

Budget: $250 – $750 USD

I'm looking for a data engineer with solid Pyspark knowledge to assist in developing a robust data storage and retrieval system, primarily focusing on a Data Warehouse.

Key Responsibilities:
- Implementing efficient data storage solutions for long-term retention and retrieval
- Ensuring data quality and validation procedures are in place
- Advising on real-time data processing capabilities

Ideal Candidate:
- Proficient in Pyspark with hands-on experience in data storage and retrieval projects
- Familiar with Data Warehousing concepts and best practices
- Able to recommend and implement appropriate real-time processing solutions
- Strong attention to detail and commitment to data quality.

Specifically, I have a Jira ticket that consists of creating an application that runs on Airflow and connects to an API with survey metadata with values such as if it was opened, if it was answered, etc, then it should generate a zip file with all the data in jsons and save it in a S3 bucket. Once in the bucket you must save the information in a Hive table.
All the code should be from Pyspark and there is a similar application that saves the raw survey data that you can take as reference.
You need to create the table and generate the code, the end client is Expedia and it should be done using their environment, I would give you the credentials and ask what you need with my colleagues.
If you like we can have a video call to explain everything better.
Related categories: Hive GitHub PySpark