Python And Scrapy for Data Extraction

Job ID: 39119815

Budget: ₹1,500 – ₹12,500 INR

For the right person with the right experience, this should be fairly straightforward. If you don’t have the right experience, it may be challenging. You must use python + scrapy.

Challenges:
1/ The site uses Imperva to block scraping.
2/ The site doesn’t (as far as I can tell) allow you to just list all the members, so you have to put in multiple locations across the UK to ensure you get all the results.

The structure of the directory doesn’t look particularly complex, and at first glance, there are three types of pages to process (listing of companies, company members, and about the company).

The expected output / deliverables are:
1/ Scraping code
2/ Scraped data i.e. structured data e.g. in json format.
3/ HTTP cache (cache of raw scraped pages)

You need to provide us with all the scraped data from the web pages in a structured format but aren’t expected to do much post-processing. However, to know that you have scraped successfully, you probably will need to do some basic processing to check you have the right number of unique records.

If you’re interested and have experience working with similar situations, we can share more information about the actual site.
Related categories: Python Web Scraping Data Mining Scrapy Data Scraping