Health Product Data Scraper
Budget: $30 – $250 USD
I need a reliable scraping solution that collects up-to-date data on health-related products sold on leading e-commerce sites. I’m primarily interested in:
• Product name and full description
• Current price (and any discounts if visible)
• Meta information such as brand, size, SKU, ASIN, etc.
• Category / sub-category path
If you already have Amazon, eBay, or Walmart parsers, that’s perfect—otherwise I’m open to whichever major marketplace you can cover first. The script must:
1. Be written in Python and use well-known libraries (Scrapy, BeautifulSoup, Requests, or Selenium—whichever fits the target site’s layout and anti-bot measures).
2. Output clean CSV or JSON that I can import directly into a spreadsheet or database.
3. Include clear setup instructions so I can run the scraper on my own machine later, tweak the search term, or change the destination file path.
4. Respect each site’s rate limits and rotate user agents / proxies where needed to avoid bans.
This is a one-off data pull right now, but if the initial run is successful I’m happy to extend the project for periodic updates or additional data points such as customer reviews and ratings.
Please outline the approach, estimated turnaround time, and any questions you have about target URLs or authentication hurdles when you bid.
• Product name and full description
• Current price (and any discounts if visible)
• Meta information such as brand, size, SKU, ASIN, etc.
• Category / sub-category path
If you already have Amazon, eBay, or Walmart parsers, that’s perfect—otherwise I’m open to whichever major marketplace you can cover first. The script must:
1. Be written in Python and use well-known libraries (Scrapy, BeautifulSoup, Requests, or Selenium—whichever fits the target site’s layout and anti-bot measures).
2. Output clean CSV or JSON that I can import directly into a spreadsheet or database.
3. Include clear setup instructions so I can run the scraper on my own machine later, tweak the search term, or change the destination file path.
4. Respect each site’s rate limits and rotate user agents / proxies where needed to avoid bans.
This is a one-off data pull right now, but if the initial run is successful I’m happy to extend the project for periodic updates or additional data points such as customer reviews and ratings.
Please outline the approach, estimated turnaround time, and any questions you have about target URLs or authentication hurdles when you bid.
Related categories:
Python
Web Scraping
Software Architecture
Data Mining
Scrapy
Data Scraping
BeautifulSoup
Selenium