AWS Data Engineer
Budget: $8 – $15 USD
I’m rolling out an AWS-native analytics platform and need a data engineer who can own our streaming and batch pipelines end-to-end. The core of the stack will revolve around Amazon EMR jobs that process events flowing in real time through Kinesis, land raw data in S3, and surface curated layers for Athena and Redshift reporting. Along the way you’ll wire in Glue for cataloging, Lambda or Step Functions for orchestration, and Lake Formation plus IAM to keep everything locked down.
What I expect you to do
• Design, build, and document scalable Spark jobs on EMR and the Kinesis streams that feed them.
• Automate everything with Terraform / CloudFormation and integrate it into an existing Git-based CI/CD pipeline.
• Instrument the workflows so I can closely track three KPIs: error rate first, performance second, and cost efficiency third.
• Tune data lake partitions, Redshift sort & distribution keys, and set up CloudWatch dashboards and alerts.
• Hand over clear runbooks and walk me through the system in a knowledge-transfer call.
This role is for someone who has already delivered EMR- plus-Kinesis solutions at scale. Show me that experience—links, metrics, or a concise write-up of what you built—and explain how you kept failure rates low while meeting performance and budget targets.
We work in two-week sprints with daily Slack check-ins, so being responsive is essential. If building bulletproof, observable data pipelines in AWS is your specialty, I’d love to see how your previous work aligns with this challenge.
What I expect you to do
• Design, build, and document scalable Spark jobs on EMR and the Kinesis streams that feed them.
• Automate everything with Terraform / CloudFormation and integrate it into an existing Git-based CI/CD pipeline.
• Instrument the workflows so I can closely track three KPIs: error rate first, performance second, and cost efficiency third.
• Tune data lake partitions, Redshift sort & distribution keys, and set up CloudWatch dashboards and alerts.
• Hand over clear runbooks and walk me through the system in a knowledge-transfer call.
This role is for someone who has already delivered EMR- plus-Kinesis solutions at scale. Show me that experience—links, metrics, or a concise write-up of what you built—and explain how you kept failure rates low while meeting performance and budget targets.
We work in two-week sprints with daily Slack check-ins, so being responsive is essential. If building bulletproof, observable data pipelines in AWS is your specialty, I’d love to see how your previous work aligns with this challenge.
Related categories:
Python
Amazon Web Services
Hadoop
Node.js
Spark
Redshift
Aws Lambda
Performance Tuning
Terraform
CI/CD