Pharm Sales Data Engineering Pipeline
Budget: ₹750 – ₹1,250 INR
Python, SQL, ETL, PySpark, Spark SQL, AWS EMR, AWS Lambda, AWS Step Functions, Amazon S3 (Data Lake – Raw & Processed Zones), AWS CloudWatch, AWS SNS, Pandas, Excel, Veeva CRM Project Overview Designed and implemented an end-to-end AWS-based data engineering pipeline to bifurcate, process, and deliver pharmaceutical sales, HCP, call activity, territory, and marketing data for Europe (EU) and Russia (RU) regions into Veeva CRM. The solution automated data ingestion from external APIs, validated and transformed high-volume datasets using Spark on EMR, and enforced multi-layer data quality checks based on business rules. Final curated datasets were delivered to Veeva CRM to support daily call planning, HCP targeting, territory alignment, and field sales insights, enabling accurate and timely decision-making for sales representatives and managers.
Roles and Responsibilities
• Built and maintained scalable medallion architecture ETL pipelines using Python, PySpark, and Spark SQL
• Implemented EU–RU data bifurcation and region-specific business logic
• Designed S3 data lake architecture (raw, processed, and error zones)
• Developed AWS Lambda triggers for schema and file validation
• Orchestrated workflows using AWS Step Functions and automated EMR jobs
• Performed data cleansing, deduplication, aggregation, and KPI generation
• Implemented data quality checks aligned with BRD requirements
• Monitored pipelines using CloudWatch logs, metrics, and SNS alerts
• Supported UAT validation and final data delivery to Veeva CRM
Roles and Responsibilities
• Built and maintained scalable medallion architecture ETL pipelines using Python, PySpark, and Spark SQL
• Implemented EU–RU data bifurcation and region-specific business logic
• Designed S3 data lake architecture (raw, processed, and error zones)
• Developed AWS Lambda triggers for schema and file validation
• Orchestrated workflows using AWS Step Functions and automated EMR jobs
• Performed data cleansing, deduplication, aggregation, and KPI generation
• Implemented data quality checks aligned with BRD requirements
• Monitored pipelines using CloudWatch logs, metrics, and SNS alerts
• Supported UAT validation and final data delivery to Veeva CRM