Comprehensive Web Scraping and Data Processing
Budget: ₹1,500 – ₹12,500 INR
Objective:
Find, standardize, and continuously update data regarding construction and infrastructure
projects and tenders in the state of California.
Part 1: Research and Data Sourcing
Task: Research and identify 5-10 reliable data sources about construction and infrastructure
projects and tenders in California.
Methodology: Use a combination of online research and language models (e.g., OpenAI's GPT
models) to identify these sources. Explicitly state how and why you used GPT or similar models
in your research process.
Part 2: Data Extraction and Standardization
Task: From the provided Table 1 and your own list, suggest methods to scrape data using
language model-based tools like OpenAI API, Mistral 7B, Llama2, or other open-source models.
Requirements:
Demonstrate how you can build data products (DPs) to scrape data from multiple sources.
Standardize the scraped data according to the guidelines provided in Table 2.
Part 3: Automation and Continuous Updating
Task: Propose a system for automating the data scraping and standardization processes.
Details:
Explain how the data sources will be continuously updated.
Describe the use of cron jobs or similar scheduling tools for ongoing data updates.
Ensure your methodology adheres to a production environment's standards.
Find, standardize, and continuously update data regarding construction and infrastructure
projects and tenders in the state of California.
Part 1: Research and Data Sourcing
Task: Research and identify 5-10 reliable data sources about construction and infrastructure
projects and tenders in California.
Methodology: Use a combination of online research and language models (e.g., OpenAI's GPT
models) to identify these sources. Explicitly state how and why you used GPT or similar models
in your research process.
Part 2: Data Extraction and Standardization
Task: From the provided Table 1 and your own list, suggest methods to scrape data using
language model-based tools like OpenAI API, Mistral 7B, Llama2, or other open-source models.
Requirements:
Demonstrate how you can build data products (DPs) to scrape data from multiple sources.
Standardize the scraped data according to the guidelines provided in Table 2.
Part 3: Automation and Continuous Updating
Task: Propose a system for automating the data scraping and standardization processes.
Details:
Explain how the data sources will be continuously updated.
Describe the use of cron jobs or similar scheduling tools for ongoing data updates.
Ensure your methodology adheres to a production environment's standards.