Healthcare Cloud Data Engineering Project
Budget: $30 – $250 USD
I’d like a hands-on, end–to-end real world data-engineering project that I can spin up in either Azure or AWS. The scenario is healthcare-focused, so the pipeline must handle claims data, patient records or survey responses from raw ingestion all the way through to an analytics-ready layer.
The build needs to show real-world thinking: automated ingestion, secure and scalable storage, and clearly documented processing logic, Data quality etc.
1) An architecture diagram that separates the data-ingestion, storage and processing layers, calling out the exact managed services used in each cloud.
2) A relational or star-schema data model.
3) Source code (Python, SQL, Spark (Databricks) or similar) that lands sample data, pre-processing data quality steps, data masking (pii/phi removal), performs transformations and writes to the final schema.
4) Documentation or instructions to do the set-up.
ONLY REAL WORLD USE CASES PLEASE. I would like to present this in an interview as my experience.
The build needs to show real-world thinking: automated ingestion, secure and scalable storage, and clearly documented processing logic, Data quality etc.
1) An architecture diagram that separates the data-ingestion, storage and processing layers, calling out the exact managed services used in each cloud.
2) A relational or star-schema data model.
3) Source code (Python, SQL, Spark (Databricks) or similar) that lands sample data, pre-processing data quality steps, data masking (pii/phi removal), performs transformations and writes to the final schema.
4) Documentation or instructions to do the set-up.
ONLY REAL WORLD USE CASES PLEASE. I would like to present this in an interview as my experience.
Related categories:
Python
Data Processing
Azure
Amazon Web Services
Data Architecture
PySpark
Data Modeling