Senior Data Engineer for Analytics
Budget: $2 – $8 USD
I need an experienced data engineer who can turn a steady stream of semi-structured information into reliable, analysis-ready datasets. The ultimate goal is data analytics, so everything you build should serve faster insight generation and easier downstream exploration.
Right now I receive semi-structured records from three key sources:
• APIs
• Logs
• Web scraping
You will design and implement the full ingestion and transformation workflow, normalising each feed, enforcing quality checks, and persisting the results in a form that scales for interactive querying. I expect you to choose tooling that makes sense—Python, Spark, Airflow, Kafka, Snowflake, Redshift, BigQuery or their equivalents—as long as the solution is robust, well-documented and cost-aware.
Deliverables must include:
• Reproducible code (version-controlled) for ingestion, parsing and transformation
• Automated tests and data quality assertions
• Deployment scripts or Terraform modules for any cloud resources you spin up
• Clear documentation describing the architecture, how to extend pipelines, and run-book style operational notes
I will consider the work complete when the pipelines run end-to-end, populate an analytics-friendly store, and a sample query proves that data from all three sources lands correctly and consistently.
Right now I receive semi-structured records from three key sources:
• APIs
• Logs
• Web scraping
You will design and implement the full ingestion and transformation workflow, normalising each feed, enforcing quality checks, and persisting the results in a form that scales for interactive querying. I expect you to choose tooling that makes sense—Python, Spark, Airflow, Kafka, Snowflake, Redshift, BigQuery or their equivalents—as long as the solution is robust, well-documented and cost-aware.
Deliverables must include:
• Reproducible code (version-controlled) for ingestion, parsing and transformation
• Automated tests and data quality assertions
• Deployment scripts or Terraform modules for any cloud resources you spin up
• Clear documentation describing the architecture, how to extend pipelines, and run-book style operational notes
I will consider the work complete when the pipelines run end-to-end, populate an analytics-friendly store, and a sample query proves that data from all three sources lands correctly and consistently.
Related categories:
Python
Data Mining
Big Data Sales
Hadoop
Redshift
Data Analytics
Apache Spark
Snowflake