Review Aggregation Tool (scrapping)
Budget: $30 – $250 USD
We are seeking a highly-skilled Python web scraping developer to create a sophisticated review aggregation tool. The primary goal is to comprehensively scrape and aggregate reviews from multiple major platforms for specific professional websites.
Primary Responsibilities:
Web Scraping Tool Creation:
Develop a tool that effectively scrapes reviews from multiple prominent online review platforms.
Implement strategies to bypass scraping restrictions without violating any terms of service or ethical guidelines.
Ensure adherence to robots.txt and employ delay strategies, changing user agents, etc., to minimize the risk of bans.
Database Management:
Design a database to store scraped reviews, associated metadata (e.g., review date, author, source), and website information (URL, active/inactive status, last scraped date).
Implement an automated mechanism to regularly update each site's reviews and recalibrate the overall ratings based on newly aggregated reviews.
Automatic detection of site activity: If a site becomes unresponsive over multiple scrapes or a set duration (e.g., 3 months), mark it as "inactive".
Automatic Discovery and Validation:
The tool should feature an automated system for discovering new professional sites related to a given industry.
Integration with popular search engine APIs or equivalent to locate potential websites.
Use Natural Language Processing (perhaps leveraging popular AI models) to validate that discovered sites truly belong to the targeted professional category.
Automated Rating System:
Design a rating mechanism that calculates an overall score for each website based on the aggregated reviews.
This score should consider weighted coefficients for reviews based on their source (i.e., certain platforms might carry a different weight compared to others).
Regular Updates & Maintenance:
The scraping tool should be set to operate at regular intervals to ensure that data remains current.
Implement checks to validate whether older sites are still active and to re-activate those that have regained activity.
Friendly Web UX Features:
Launch scraping, maintenance, discover a new site from a category/sector using popular search engine APIs.
Add sites manually or using a CSV file, access to the database via the UX, and change the coefficients for each source regarding site rating.
Required Skills and Experience:
Profound expertise in Python and associated web scraping frameworks such as popular Python scraping libraries.
Strong SQL database knowledge and experience with database optimization.
Familiarity with NLP for content validation.
Previous experience in circumventing scraping limitations without breaching ethical or legal standards.
Demonstrated ability to automate complex processes.
Additional Notes:
Please provide examples of similar projects you've undertaken.
Primary Responsibilities:
Web Scraping Tool Creation:
Develop a tool that effectively scrapes reviews from multiple prominent online review platforms.
Implement strategies to bypass scraping restrictions without violating any terms of service or ethical guidelines.
Ensure adherence to robots.txt and employ delay strategies, changing user agents, etc., to minimize the risk of bans.
Database Management:
Design a database to store scraped reviews, associated metadata (e.g., review date, author, source), and website information (URL, active/inactive status, last scraped date).
Implement an automated mechanism to regularly update each site's reviews and recalibrate the overall ratings based on newly aggregated reviews.
Automatic detection of site activity: If a site becomes unresponsive over multiple scrapes or a set duration (e.g., 3 months), mark it as "inactive".
Automatic Discovery and Validation:
The tool should feature an automated system for discovering new professional sites related to a given industry.
Integration with popular search engine APIs or equivalent to locate potential websites.
Use Natural Language Processing (perhaps leveraging popular AI models) to validate that discovered sites truly belong to the targeted professional category.
Automated Rating System:
Design a rating mechanism that calculates an overall score for each website based on the aggregated reviews.
This score should consider weighted coefficients for reviews based on their source (i.e., certain platforms might carry a different weight compared to others).
Regular Updates & Maintenance:
The scraping tool should be set to operate at regular intervals to ensure that data remains current.
Implement checks to validate whether older sites are still active and to re-activate those that have regained activity.
Friendly Web UX Features:
Launch scraping, maintenance, discover a new site from a category/sector using popular search engine APIs.
Add sites manually or using a CSV file, access to the database via the UX, and change the coefficients for each source regarding site rating.
Required Skills and Experience:
Profound expertise in Python and associated web scraping frameworks such as popular Python scraping libraries.
Strong SQL database knowledge and experience with database optimization.
Familiarity with NLP for content validation.
Previous experience in circumventing scraping limitations without breaching ethical or legal standards.
Demonstrated ability to automate complex processes.
Additional Notes:
Please provide examples of similar projects you've undertaken.