Data Lakehouse Setup with Apache Iceberg, Nessie, MinIO, and Dremio on Kubernetes

Job ID: 38504011

Budget: ₹1,500 – ₹12,500 INR

We are looking for a skilled data engineer to design and implement a data lakehouse architecture using Apache Iceberg, Project Nessie, MinIO, and Dremio, with deployment on Kubernetes. This project includes providing Python functions for common data operations, ensuring they are well-documented for ease of use.

Scope of Work:

1. Data Lakehouse Architecture:

- Deploy a data lakehouse using Apache Iceberg for table management and Project Nessie for catalog and versioning.
- Utilize MinIO as the S3-compatible storage backend.
- Implement Dremio as the platform for managing, querying, and optimizing the lakehouse.

2. Kubernetes Deployment:

- Deploy the entire setup on Kubernetes, ensuring scalability and reliability.
- Provide a Helm chart for simplified deployment and configuration.
- Include a step-by-step installation guide for setting up Kubernetes, Helm, and the data lakehouse from scratch.

3. Python Functions:

- **Table Operations:**
- Function for creating new tables with partitions and indices in Iceberg.
- Function for modifying the schema of existing tables.
- **Data Ingestion:**
- Function to write data from a CSV file to the lakehouse.
- Function to write data from an in-memory pandas DataFrame.
- Function to write data from a CSV file loaded in memory.
- **Data Retrieval:**
- Function to load data from the lakehouse based on specified search criteria.

Requirements:

- **Expertise in Apache Iceberg and Project Nessie.**
- **Experience with Kubernetes and Helm for container orchestration and deployment.**
- **Proficiency in Python, particularly in data engineering tasks such as pandas and CSV handling.**
- **Experience with object storage solutions like MinIO or Amazon S3.**
- **Experience with Dremio or similar data lake management tools.**

**Documentation:**

- Detailed instructions for setting up the Kubernetes environment, Helm, and the lakehouse.
- Clear, well-commented documentation for each Python function, making it accessible for both developers and data engineers.
- A user guide on how to use the Python functions for data ingestion and retrieval.

Project Deliverables:

- A fully functional data lakehouse running on Kubernetes.
- A Helm chart for easy deployment.
- Python functions with comprehensive documentation for managing, ingesting, and retrieving data from the lakehouse.
- Step-by-step documentation for setup and usage, aimed at beginners.

Timeline:
Please provide your estimated timeline for completing this project, including milestones for initial setup, deploying data lakehouse, Python function development, and final documentation delivery.