Automated Web Scraping System for Event Info
Budget: $10 – $30 USD
The development of an automated web scraping system is required to extract event information from multiple websites. The system must process data from a list of 34 venues and bars, whose information is contained in a provided Excel file. The goal is to automatically generate a weekly event sheet in Google Sheets.
The developer will be responsible for:
- Conducting a thorough technical analysis of each of the 34 sources to determine their scraping feasibility and the most appropriate method. It is anticipated that not all sources will have an optimal structure for automated scraping.
- Implementing scraping for all technically feasible sources.
- Managing the heterogeneity of the websites, including proprietary websites, WordPress-based sites, and ticketing platforms.
- Developing solutions for scraping dynamic content (generated with JavaScript) from sources that require it.
It is important to note that Instagram scraping is not included in the scope of this phase of the project.
The data to be extracted for each event, when available, includes:
- Event Name
- Date
- Time
- Venue/Bar
- Ticket Price
- Description
- Performing Musicians (first and last name separated by commas in a single cell)
- Link to the event
- Source URL
Technical and Functional Requirements:
- Standardized formats for dates (yyyy-mm-dd) and times (hh:mm).
- Implementation of a basic mechanism for duplicate removal, based on the combination of event name, date, and venue.
- Automatic generation of a weekly event sheet in Google Sheets, covering the period from Monday to Sunday.
- Use of the Google Sheets API for interaction with the spreadsheets.
- The system must be implemented and run on a Google Cloud (GCP) virtual machine, configured to allow manual on/off operation by the client.
- The source code must be documented, modular, and designed to be scalable. Clear and concise documentation on the use and maintenance of the system must be provided.
The proposal should include information on the technologies to be used, the proposed methodology for evaluating the feasibility of scraping, prior experience with complex scraping projects, and an estimate of maintenance time and costs.
The developer will be responsible for:
- Conducting a thorough technical analysis of each of the 34 sources to determine their scraping feasibility and the most appropriate method. It is anticipated that not all sources will have an optimal structure for automated scraping.
- Implementing scraping for all technically feasible sources.
- Managing the heterogeneity of the websites, including proprietary websites, WordPress-based sites, and ticketing platforms.
- Developing solutions for scraping dynamic content (generated with JavaScript) from sources that require it.
It is important to note that Instagram scraping is not included in the scope of this phase of the project.
The data to be extracted for each event, when available, includes:
- Event Name
- Date
- Time
- Venue/Bar
- Ticket Price
- Description
- Performing Musicians (first and last name separated by commas in a single cell)
- Link to the event
- Source URL
Technical and Functional Requirements:
- Standardized formats for dates (yyyy-mm-dd) and times (hh:mm).
- Implementation of a basic mechanism for duplicate removal, based on the combination of event name, date, and venue.
- Automatic generation of a weekly event sheet in Google Sheets, covering the period from Monday to Sunday.
- Use of the Google Sheets API for interaction with the spreadsheets.
- The system must be implemented and run on a Google Cloud (GCP) virtual machine, configured to allow manual on/off operation by the client.
- The source code must be documented, modular, and designed to be scalable. Clear and concise documentation on the use and maintenance of the system must be provided.
The proposal should include information on the technologies to be used, the proposed methodology for evaluating the feasibility of scraping, prior experience with complex scraping projects, and an estimate of maintenance time and costs.
Related categories:
Python
Data Processing
Web Scraping
Software Architecture
Data Mining
Data Extraction
Automation
Data Management