Senior Data Engineer

Job ID: 40200033

Budget: $25 – $50 USD

Job Description:

We are looking for a Senior Data Engineer with expertise in data engineering, ETL processes, cloud platforms, and data warehousing. You will design, build, and maintain scalable and efficient data pipelines that drive our data operations across multiple environments. The ideal candidate will have hands-on experience with cloud platforms (AWS, Azure), real-time streaming technologies, data quality frameworks, and data warehousing solutions (Redshift, Snowflake, etc.). You will work collaboratively with data scientists, analysts, and cross-functional teams to ensure the smooth flow of data across the organization.

Key Responsibilities:

- Develop and maintain end-to-end ETL pipelines, handling large volumes of structured and unstructured data.

- Design and implement data models in Amazon Redshift, Snowflake, Azure SQL, and PostgreSQL to enable scalable, performant analytics.

- Leverage Apache Airflow and Azure Data Factory for orchestration and scheduling of complex data workflows.

- Build and optimize real-time data streams using Apache Kafka, AWS Kinesis, and Kinesis Firehose.

- Collaborate with teams to ensure data quality and reliability through automated testing using Great Expectations and PyTest.

- Utilize cloud platforms such as AWS (EC2, S3, Lambda, RDS, Glue) and Microsoft Azure (Blob Storage, Functions, VMs) for data ingestion, processing, and storage.

- Ensure smooth integration with Power BI, Tableau, and other BI tools, providing teams with clean, actionable data for decision-making.

- Troubleshoot and optimize performance across data pipelines, improving query performance, data latency, and cost efficiency.

- Continuously improve data architecture, ensuring robust data governance and security practices.

- Maintain clear, well-documented workflows, data models, and pipelines to ensure team collaboration and future scalability.

Required Skills and Qualifications:

- 6+ years of experience in data engineering and hands-on ETL development.

- Proficiency in Python, SQL, Java, Scala, Bash, and PowerShell.

- Strong experience with data orchestration and scheduling tools like Apache Airflow, Azure Data Factory, and CloudWatch Events.

- Expertise in cloud data platforms (AWS, Azure), with hands-on experience in AWS Glue, S3, Lambda, and RDS.

- In-depth knowledge of data warehousing (e.g., Amazon Redshift, Snowflake, Azure SQL, PostgreSQL).

- Experience with real-time data streaming technologies, including Apache Kafka, AWS Kinesis, and Kinesis Firehose.

- Familiarity with data quality tools like Great Expectations and PyTest.

- Strong experience with BI and data visualization tools (e.g., Power BI, Tableau, SSRS, Power Query).

- Solid understanding of data modeling, incremental loading, data partitioning, and query optimization.

- Strong communication skills with the ability to collaborate with technical and non-technical teams.

Preferred Qualifications:

- Experience with machine learning pipelines or AI tools is a plus.

- Familiarity with Microsoft Fabric is a plus.

- Experience in large-scale system implementations or complex data projects.

Why Join Us?

- Collaborative culture focused on innovation and growth.

- Work with cutting-edge technologies and gain hands-on experience with modern data platforms.

- Remote-first environment with flexibility to work from anywhere.

- Competitive compensation and benefits package.