Automated Web Scraping System Development
Budget: $250 – $750 USD
Position Description - Web Scraping Developer
Project Objective
Develop and maintain an automated Web Scraping system to collect product data from marketplaces, including information on prices, sales history, availability and images. The goal is to create an up-to-date database with detailed data on items, ranging from new to used.
Work Scope
1. Data Collection
Automatic extraction of information on products, prices, offers and item details.
Collection of data related to new and used items.
Organize and store information efficiently.
2. Update and Synchronization
Implementation of a system to run daily queries, keeping information always up to date.
Management of price and availability changes with notifications for updates.
3. Proxy and Anti-Blocking
Use of rotating proxies to avoid blocking during data collection.
Application of random delays and obfuscation techniques.
4. Storage and Structuring
Storage of collected data in a structured database.
Ensure queryability and scalability of information.
Implementation of API for data query.
5. Monitoring and Logs
Monitoring to ensure continuous system operation.
Implementation of detailed logs to track errors or failures during data collection.
Technical Requirements
Programming Language: Python
Libraries: Scrapy, BeautifulSoup, Selenium
Database: PostgreSQL / Firebase / MongoDB
Proxies: Proxy tools such as BrightData or equivalent
Automation: Use of Cron Jobs or Celery to schedule tasks.
Project Objective
Develop and maintain an automated Web Scraping system to collect product data from marketplaces, including information on prices, sales history, availability and images. The goal is to create an up-to-date database with detailed data on items, ranging from new to used.
Work Scope
1. Data Collection
Automatic extraction of information on products, prices, offers and item details.
Collection of data related to new and used items.
Organize and store information efficiently.
2. Update and Synchronization
Implementation of a system to run daily queries, keeping information always up to date.
Management of price and availability changes with notifications for updates.
3. Proxy and Anti-Blocking
Use of rotating proxies to avoid blocking during data collection.
Application of random delays and obfuscation techniques.
4. Storage and Structuring
Storage of collected data in a structured database.
Ensure queryability and scalability of information.
Implementation of API for data query.
5. Monitoring and Logs
Monitoring to ensure continuous system operation.
Implementation of detailed logs to track errors or failures during data collection.
Technical Requirements
Programming Language: Python
Libraries: Scrapy, BeautifulSoup, Selenium
Database: PostgreSQL / Firebase / MongoDB
Proxies: Proxy tools such as BrightData or equivalent
Automation: Use of Cron Jobs or Celery to schedule tasks.
Related categories:
PHP
Business, Accounting, Human Resources & Legal
Web Scraping
Software Architecture
Data Mining