AWS Glue: Incremental Fetch from BigQuery and Store on S3
Budget: $30 – $250 USD
Seeking an AWS Glue expert to assist with fetching analytical data from Google BigQuery and storing it on S3 in Parquet or CSV format. The job includes setting up an incremental data extraction process that runs daily.
Current Status:
Query: Already prepared.
Connection: The connection to BigQuery is set up and ready from within Glue Studio.
Challenge:
1- I need assistance configuring Glue ( Glue Notebook or visual Job) to handle date-partitioned tables in BigQuery and load data from there incrementally.
2- configure S3 crawler to scan the bucket and push new data to DB
Background: I previously implemented this workflow using QlikView script, but I am now transitioning to AWS. Looking for guidance on best practices specific to AWS.
Ideal candidates should have:
- Extensive experience with AWS Glue, AWS Glue NoteBook (PySpark) and S3 crawlers
Current Status:
Query: Already prepared.
Connection: The connection to BigQuery is set up and ready from within Glue Studio.
Challenge:
1- I need assistance configuring Glue ( Glue Notebook or visual Job) to handle date-partitioned tables in BigQuery and load data from there incrementally.
2- configure S3 crawler to scan the bucket and push new data to DB
Background: I previously implemented this workflow using QlikView script, but I am now transitioning to AWS. Looking for guidance on best practices specific to AWS.
Ideal candidates should have:
- Extensive experience with AWS Glue, AWS Glue NoteBook (PySpark) and S3 crawlers