Build BigQuery Data Pipelines
Budget: ₹400 – ₹750 INR
I’m putting together a set of production-grade data pipelines on Google Cloud Platform and would like an experienced hand to own the build. The job is centred on pipeline development rather than analysis or migration tasks.
Scope
• Source systems: relational databases we already run in GCP plus batches of parquet and CSV files landing in Cloud Storage.
• Destination: curated, partitioned tables in BigQuery with appropriate clustering and cost-efficient storage settings.
Work I need from you
– Design the end-to-end flow, including ingestion, transformation, and load steps.
– Write modular, reusable code or SQL that can be version-controlled and promoted across environments.
– Automate orchestration (Cloud Composer, Cloud Functions, or another native GCP option—open to your recommendation).
– Implement monitoring, alerting, and simple rollback or re-run logic so failures are surfaced quickly.
– Provide a concise hand-over note explaining how to extend or tweak the pipeline.
Acceptance criteria
1. A repeatable deployment (scripts or Terraform) that spins up all required GCP resources.
2. Successful extraction from both the database and Cloud Storage sample sets, transformed and loaded into BigQuery.
3. Query results in BigQuery match source record counts and basic data-quality checks.
4. Logging visible in Cloud Logging and alerts routed to our Slack channel.
If you’ve previously delivered similar BigQuery-focused pipelines and can start soon, let’s talk details.
Scope
• Source systems: relational databases we already run in GCP plus batches of parquet and CSV files landing in Cloud Storage.
• Destination: curated, partitioned tables in BigQuery with appropriate clustering and cost-efficient storage settings.
Work I need from you
– Design the end-to-end flow, including ingestion, transformation, and load steps.
– Write modular, reusable code or SQL that can be version-controlled and promoted across environments.
– Automate orchestration (Cloud Composer, Cloud Functions, or another native GCP option—open to your recommendation).
– Implement monitoring, alerting, and simple rollback or re-run logic so failures are surfaced quickly.
– Provide a concise hand-over note explaining how to extend or tweak the pipeline.
Acceptance criteria
1. A repeatable deployment (scripts or Terraform) that spins up all required GCP resources.
2. Successful extraction from both the database and Cloud Storage sample sets, transformed and loaded into BigQuery.
3. Query results in BigQuery match source record counts and basic data-quality checks.
4. Logging visible in Cloud Logging and alerts routed to our Slack channel.
If you’ve previously delivered similar BigQuery-focused pipelines and can start soon, let’s talk details.
Related categories:
SQL
MySQL
Hadoop
PostgreSQL
Data Warehousing
Google Cloud Platform
PySpark
BigQuery