Scrapy Scraping and Web Crawling parallel execution
Budget: $250 – $750 USD
Hey,
I need a web crawler for a certain website that discovers links by itself and manages to intelligent scrape some content from it.
Requirements:
- Can run multiple instances at the same time (so 2/3 crawler for the same website at the same time)
- Prevents the crawling of a link twice within a certain period (crawl a link only every 2 weeks)
- Crawl (find new links)
- Extract around 10 data points
- Save to database
- Execute JavaScript
- Use Proxies (provided by me)
- Has restriction of crawls per minute
- Uses state-of-the-art methods for hiding that it is not a person (browser fingerprint etc.)
Please state your experience and also how you think you can solve the need to manage the multiple instances and prevent crawling twice.
Thnaks.
I need a web crawler for a certain website that discovers links by itself and manages to intelligent scrape some content from it.
Requirements:
- Can run multiple instances at the same time (so 2/3 crawler for the same website at the same time)
- Prevents the crawling of a link twice within a certain period (crawl a link only every 2 weeks)
- Crawl (find new links)
- Extract around 10 data points
- Save to database
- Execute JavaScript
- Use Proxies (provided by me)
- Has restriction of crawls per minute
- Uses state-of-the-art methods for hiding that it is not a person (browser fingerprint etc.)
Please state your experience and also how you think you can solve the need to manage the multiple instances and prevent crawling twice.
Thnaks.