web crawler
Budget: ₹12,500 – ₹37,500 INR
I am looking to create a web crawler that will extract specific data from e-commerce stores online. The goal is to collect text content, which needs to be cleaned and formatted so that it is ready to be used. I require a solution that will extract relevant information quickly and accurately - I will not accept anything less. If you think you have the skills necessary to deliver a successful project, then I invite you to submit a bid for this project. Thank you for your time!
This will require someone to
1. Crawl a parent page
2. Get all the links in that page
3. Hit that link and then get few data from the page
4. Create an excel with that data
Things to note ,
1.There should be some sleep Time introduced if there is continous hits the IP address can be blocked due to rate Limiting
2. Need a MultiThreaded application which can run parallel crawls, to make sure the crawling doesnt take hours ( 1 crawl could contain 1000+ links to hit)
3. If Program crashes/Needs to be stopped the previously crawled data should not be lost - One solution is intermediate writes
4. Open to run on local machine or cloud ( Ready to invest more if cloud can help scale)
This will require someone to
1. Crawl a parent page
2. Get all the links in that page
3. Hit that link and then get few data from the page
4. Create an excel with that data
Things to note ,
1.There should be some sleep Time introduced if there is continous hits the IP address can be blocked due to rate Limiting
2. Need a MultiThreaded application which can run parallel crawls, to make sure the crawling doesnt take hours ( 1 crawl could contain 1000+ links to hit)
3. If Program crashes/Needs to be stopped the previously crawled data should not be lost - One solution is intermediate writes
4. Open to run on local machine or cloud ( Ready to invest more if cloud can help scale)