Build Robust Database Pipelines
Budget: ₹12,500 – ₹37,500 INR
I need an experienced data engineer to design and implement fully automated data pipelines that pull from our production databases and deliver clean, query-ready tables to our analytics environment. The work is focused on Data Pipelines only—no warehousing architecture or ad-hoc analysis is required.
Here is the scope I have in mind:
• Connect to multiple relational databases (currently MySQL and PostgreSQL, with the possibility of others) and set up reliable, incremental extractions.
• Transform and load the data so that downstream analysts can query consistent, well-documented tables.
• Embed resilience features such as retry logic, data-quality checks, and schema-change handling.
• Orchestrate everything on a modern scheduling platform (Airflow or an equivalent) and write the code primarily in Python and SQL.
• Provide clear documentation plus a hand-over session so my in-house team can maintain and extend the pipelines.
Acceptance criteria for the final hand-off:
1. End-to-end run completes without errors against a staging database.
2. At least 95 % of rows pass data-quality checks defined together.
3. Configuration and credentials are externalised; no secrets hard-coded.
4. README explains setup, deployment steps, and how to add new tables.
5. All code stored in our private Git repository with meaningful commit history.
If you have a track record of building production-grade ETL / ELT workflows from database sources and can start soon, I’d love to review your approach and timeline.
Here is the scope I have in mind:
• Connect to multiple relational databases (currently MySQL and PostgreSQL, with the possibility of others) and set up reliable, incremental extractions.
• Transform and load the data so that downstream analysts can query consistent, well-documented tables.
• Embed resilience features such as retry logic, data-quality checks, and schema-change handling.
• Orchestrate everything on a modern scheduling platform (Airflow or an equivalent) and write the code primarily in Python and SQL.
• Provide clear documentation plus a hand-over session so my in-house team can maintain and extend the pipelines.
Acceptance criteria for the final hand-off:
1. End-to-end run completes without errors against a staging database.
2. At least 95 % of rows pass data-quality checks defined together.
3. Configuration and credentials are externalised; no secrets hard-coded.
4. README explains setup, deployment steps, and how to add new tables.
5. All code stored in our private Git repository with meaningful commit history.
If you have a track record of building production-grade ETL / ELT workflows from database sources and can start soon, I’d love to review your approach and timeline.
Related categories:
Python
Data Processing
SQL
MySQL
Database Programming
Data Extraction
Data Integration
ETL