Lead GCP Data engineer
Budget: ₹12,500 – ₹37,500 INR
Key Responsibilities: experience 14+Years
● Design, develop, test, and maintain scalable ETL data pipelines using Python.
● Architect the enterprise solutions with various technologies like Kafka, multi-cloud services, auto-scaling using GKE, Load balancers, APIGEE proxy API management, DBT, using LLMs as needed in the solution, redaction of sensitive information, DLP (Data Loss Prevention) etc.
● Work extensively on Google Cloud Platform (GCP) services such as:
○ Data-flow for real-time and batch data processing
○ Cloud Functions for lightweight serverless compute
○ BigQuery for data warehousing and analytics
○ Cloud Composer for orchestration of data workflows (on Apache Airflow)
○ Google Cloud Storage (GCS) for managing data at scale
○ IAM for access control and security
○ Cloud Run for containerized applications Should have experience in the following areas :
○ API framework: Python FastAPI
○ Processing engine: Apache Spark
○ Messaging and streaming data processing: Kafka
○ Storage: MongoDB, Redis/Big table
○ Orchestration: Airflow
○ Experience in deployments in GKE, Cloud Run.
● Perform data ingestion from various sources and apply transformation and cleansing logic to ensure high-quality data delivery.
● Implement and enforc
● Design, develop, test, and maintain scalable ETL data pipelines using Python.
● Architect the enterprise solutions with various technologies like Kafka, multi-cloud services, auto-scaling using GKE, Load balancers, APIGEE proxy API management, DBT, using LLMs as needed in the solution, redaction of sensitive information, DLP (Data Loss Prevention) etc.
● Work extensively on Google Cloud Platform (GCP) services such as:
○ Data-flow for real-time and batch data processing
○ Cloud Functions for lightweight serverless compute
○ BigQuery for data warehousing and analytics
○ Cloud Composer for orchestration of data workflows (on Apache Airflow)
○ Google Cloud Storage (GCS) for managing data at scale
○ IAM for access control and security
○ Cloud Run for containerized applications Should have experience in the following areas :
○ API framework: Python FastAPI
○ Processing engine: Apache Spark
○ Messaging and streaming data processing: Kafka
○ Storage: MongoDB, Redis/Big table
○ Orchestration: Airflow
○ Experience in deployments in GKE, Cloud Run.
● Perform data ingestion from various sources and apply transformation and cleansing logic to ensure high-quality data delivery.
● Implement and enforc
Related categories:
Python
Data Processing
Hadoop
Map Reduce
Google Cloud Platform
ETL
BigQuery
Apache Kafka
Apache Spark
FastAPI