Advanced Web crawling/Web automation! -- 2
Budget: €2 – €10 EUR
For one of our projects, we scrape from different car marketplaces (e.g., https://www.pkw.de/) the URL link of the car ads (URL link = https://suche.pkw.de/fahrzeuge/details/372161012371653506).
We use costume python (Beautiful Soup) spiders (see pkw.py). Our spiders do the job but scrape just a part of the URLs.
Reason:
The car marketplaces (e.g., https://www.pkw.de/) use pagination limitations. The spiders use page filters to get all pages. The issue is that the filters are not chosen very elegantly and we miss a lot of data.
The task is to find better filter combinations or/and choose the right start Url to get all (! important 100%) URLs.
It is important that the spiders are very fast because some of the websites have a huge Nr. of ads.
Therefore, an elegant selection of filter combinations is required, to be fast and avoid unnecessary spider running for empty filters. To optimize this part, deep learning can be used to move to the next filter, when all ads are scraped in the past filter.
Deep learning skills would be helpful, but not required.
Please only advanced developers, with experience in this field – otherwise we both waste some hours.
We are open to discussing faster or any other elegant solution.
We are very flexible for any type of payment (fixed price, weekly, monthly, etc.).
We use a complex tool and search for talents for long-term coop.
We would love to invite you to our team if it is the right project for you
We use costume python (Beautiful Soup) spiders (see pkw.py). Our spiders do the job but scrape just a part of the URLs.
Reason:
The car marketplaces (e.g., https://www.pkw.de/) use pagination limitations. The spiders use page filters to get all pages. The issue is that the filters are not chosen very elegantly and we miss a lot of data.
The task is to find better filter combinations or/and choose the right start Url to get all (! important 100%) URLs.
It is important that the spiders are very fast because some of the websites have a huge Nr. of ads.
Therefore, an elegant selection of filter combinations is required, to be fast and avoid unnecessary spider running for empty filters. To optimize this part, deep learning can be used to move to the next filter, when all ads are scraped in the past filter.
Deep learning skills would be helpful, but not required.
Please only advanced developers, with experience in this field – otherwise we both waste some hours.
We are open to discussing faster or any other elegant solution.
We are very flexible for any type of payment (fixed price, weekly, monthly, etc.).
We use a complex tool and search for talents for long-term coop.
We would love to invite you to our team if it is the right project for you