AI-Driven Product Data Scraper
Budget: $15 – $25 USD
I need a reliable script that can automatically pull detailed product information from several well-known e-commerce sites that I will specify once we begin. The sole purpose is clean, repeatable data extraction, not analytics or price comparisons.
Here’s what matters most to me:
• Accuracy and completeness of fields such as title, price, images, description, rating, and availability.
• Resilience to site changes and anti-bot measures (dynamic content, pagination, CAPTCHAs).
• Output delivered in structured CSV or JSON so I can drop it straight into my own pipeline.
• Clear instructions on how to run or schedule the scraper on a standard Python environment; feel free to use Scrapy, BeautifulSoup, Selenium, Playwright, or any modern AI/ML technique that helps maintain extraction quality.
I’ll provide the exact URLs and the final list of fields after we agree on an approach. If you’ve built scrapers for Amazon, eBay, Walmart, or similar platforms, share a quick example or demo URL so I can see your style.
I consider the job done when I can run the code on my machine, point it at the target URLs, and receive a correctly formatted dataset with no missing columns.
Here’s what matters most to me:
• Accuracy and completeness of fields such as title, price, images, description, rating, and availability.
• Resilience to site changes and anti-bot measures (dynamic content, pagination, CAPTCHAs).
• Output delivered in structured CSV or JSON so I can drop it straight into my own pipeline.
• Clear instructions on how to run or schedule the scraper on a standard Python environment; feel free to use Scrapy, BeautifulSoup, Selenium, Playwright, or any modern AI/ML technique that helps maintain extraction quality.
I’ll provide the exact URLs and the final list of fields after we agree on an approach. If you’ve built scrapers for Amazon, eBay, Walmart, or similar platforms, share a quick example or demo URL so I can see your style.
I consider the job done when I can run the code on my machine, point it at the target URLs, and receive a correctly formatted dataset with no missing columns.