Python, Spark ETL Developer
Budget: $20 – $30 NZD
I am looking for a seasoned Python, Spark developer to assist with ETL processes who has worked on hash based algorithms to solve the given problem.
Ideal Skills:
- Proficiency in Python ,Spark and SQL is a must.
- Previous experience in data extraction, transformation, and loading.
I have 1000 large tables in the source SQL Server database where DML operations are performed. We have a nightly job that processes these tables, but it currently fails to capture deletes from the source tables. The destination is a Snowflake database. However, I face several constraints:
a) I cannot make any changes to the source tables and databases.
b) I cannot use Change Data Capture (CDC).
c) A full table refresh is not feasible due to the large size of the source tables.
Given these constraints, what solutions would you recommend to ensure deletes are accurately reflected in the Snowflake database?"
Data Source:
- The data will be primarily sourced from databases, requiring a good understanding of database systems and SQL.
Data Volume:
- The anticipated volume of data to be processed is between 1 TB to 10 TB.
Ideal Skills and Experience:
- Proficient in Python and Spark with a strong understanding of ETL processes.
- Experience working with database systems and SQL.
- Capability to handle a data volume ranging from 1 TB to 10 TB.
The successful candidate will be able to efficiently manage the ETL processes, ensuring smooth data transformation and loading, allowing us to leverage our database effectively. Let's discuss the project further!
Ideal Skills:
- Proficiency in Python ,Spark and SQL is a must.
- Previous experience in data extraction, transformation, and loading.
I have 1000 large tables in the source SQL Server database where DML operations are performed. We have a nightly job that processes these tables, but it currently fails to capture deletes from the source tables. The destination is a Snowflake database. However, I face several constraints:
a) I cannot make any changes to the source tables and databases.
b) I cannot use Change Data Capture (CDC).
c) A full table refresh is not feasible due to the large size of the source tables.
Given these constraints, what solutions would you recommend to ensure deletes are accurately reflected in the Snowflake database?"
Data Source:
- The data will be primarily sourced from databases, requiring a good understanding of database systems and SQL.
Data Volume:
- The anticipated volume of data to be processed is between 1 TB to 10 TB.
Ideal Skills and Experience:
- Proficient in Python and Spark with a strong understanding of ETL processes.
- Experience working with database systems and SQL.
- Capability to handle a data volume ranging from 1 TB to 10 TB.
The successful candidate will be able to efficiently manage the ETL processes, ensuring smooth data transformation and loading, allowing us to leverage our database effectively. Let's discuss the project further!