Java Kafka EMR Hadoop Data Specialist
Budget: ₹600 – ₹1,500 INR
I’m building a data-processing and analytics platform on AWS EMR that pulls events from Kafka, lands them in Hadoop, and serves refined datasets to downstream teams. I need an engineer who can wire everything together, tune it for performance, and leave me with maintainable jobs and clear documentation.
The stack you’ll be working with revolves around:
• Apache Spark for large-scale transformations
• Hive for ad-hoc SQL and schema management
• HBase for low-latency look-ups
Your responsibilities include designing the end-to-end flow from Kafka topics into EMR, choosing optimal storage formats, writing Spark jobs, defining Hive schemas, and configuring HBase tables. I’ll rely on you for best-practice partitioning, security hardening, and job orchestration using the tools you know best.
Deliverables
1. Infrastructure scripts (CloudFormation, Terraform, or similar) that spin up the EMR environment
2. Reproducible Spark jobs stored in Git
3. Hive tables with sample queries
4. Operational HBase setup and example API calls
5. A concise runbook covering deployment, monitoring, and troubleshooting
When you apply, show me past work that proves you’ve handled similar Kafka → EMR → Spark/Hive/HBase pipelines at scale. Screenshots, code snippets, or short case studies are perfect; no lengthy proposal needed.
I’d like to see a first working pipeline within two weeks, so let me know your availability and any questions you might have.
with AWS experience is plus
The stack you’ll be working with revolves around:
• Apache Spark for large-scale transformations
• Hive for ad-hoc SQL and schema management
• HBase for low-latency look-ups
Your responsibilities include designing the end-to-end flow from Kafka topics into EMR, choosing optimal storage formats, writing Spark jobs, defining Hive schemas, and configuring HBase tables. I’ll rely on you for best-practice partitioning, security hardening, and job orchestration using the tools you know best.
Deliverables
1. Infrastructure scripts (CloudFormation, Terraform, or similar) that spin up the EMR environment
2. Reproducible Spark jobs stored in Git
3. Hive tables with sample queries
4. Operational HBase setup and example API calls
5. A concise runbook covering deployment, monitoring, and troubleshooting
When you apply, show me past work that proves you’ve handled similar Kafka → EMR → Spark/Hive/HBase pipelines at scale. Screenshots, code snippets, or short case studies are perfect; no lengthy proposal needed.
I’d like to see a first working pipeline within two weeks, so let me know your availability and any questions you might have.
with AWS experience is plus
Related categories:
Data Processing
Database Administration
Big Data Sales
Hadoop
Map Reduce
Scala
R Programming Language
Hive
HBase
Apache Spark