Website scraper - to review and finalize an ongoing project

Job ID: 37431337

Budget: $250 – $750 USD

Website Scraper - Review and Finalize Ongoing Project

Purpose:
- Data extraction from a complex website

Skills and Experience:
- Strong experience in web scraping
- Proficient in Python or other relevant programming languages
- Familiarity with complex website structures and handling dynamic content
- Knowledge of data extraction techniques and tools
- Amazon Web Services

Output Format:
- Extracted data should be saved in a JSON file

Project Details:
- We are looking for a skilled website scraper to review and finalize an ongoing project.
- The purpose of the website scraper is to extract data from a complex website.
- The preferred output format for the extracted data is a JSON file.
- The website to be scraped is of a complex nature, requiring expertise in handling dynamic content and complex website structures.
- The selected freelancer should have a strong background in web scraping and proficiency in relevant programming languages, such as Python.
- The ideal candidate should also be familiar with various data extraction techniques and tools.
- Attention to detail and accuracy are crucial for this project.

We have a mostly developed scraping project – the trickiest parts are ready I think - which needs to be reviewed and finalized due to the unforeseen absence of the original coder. We were very satisfied with what he delivered till a point where he just vanished and we – or the Freelancer.com support staff – cannot reach him since weeks.



The original task was: (1.) to identify a total of 350 products pre-selected by the client on the websites of nine Internet data sources (Hungarian-language online stores of retail chains), some of which are offered by several retailers and some of which are sold by only one retailer.
(2.) The parameters of the identified products have to be provided in a structured format to the background database manager.
(3.) The data updates should be synchronized with the update time of the relevant data source, in terms of price and availability of the product, probably on a daily basis, in order to ensure that the offers are as accurate as possible. (4.) The data should be updated on a daily basis over a two-week testing period to ensure accuracy (apart from products that become unavailable during the day due to stock depletion.)
Based on the product URLs, the following product data should be provided to the backend database: - Product name - Product description - Product image - Product packaging (unit and quantity) - Product price In the case of a promotional price, the start and end of the promotion If available, also the non-special price - Other product attributes (ingredients, vegan not vegan, etc.) - Product unique identifier (if any)
(5.) The implementation should not give the impression of a "bot" to avoid possible flagging, e.g.: properly configured "real life" web client (user-agent, etc.), proxy configuration from list (if you want to run the scraper from multiple IP addresses).
(6.) It would be an advantage to develop a solution that can efficiently handle product data for the entire portfolio of the given sources in the future, if the test is successful (up to 60-70E products per source).
The sponsor considered a timeframe of 2 weeks as sufficient to complete the task.

The programming language and framework were optional and finally the freelancer chose Phyton/AWS lambda.


What has been achieved?

The data types have been identified and the scrapers have been prepared for the individual data sources (urls), which are able to return the appropriate information in the required format for the background database.

What is ahead of us?

In short it's step # 3-4-5 by taking into consideration the advantage described in #6:

The algorithm must be prepared, which provides the scrapers with the urls of the given website that we have predefined in a list, so that we can automatically read the current data of all products. (After the first identification, it is expected that only the product price and warehouse availability status will change from the given url.)

Data from different retail websites must be standardized so that they can be compared with each other.

The product data must be automatically updated using Amazon Web Services and transferred to the background database on a daily basis.

When the automatic update is ready, it is necessary to determine by testing which data source, when during the day it is worth performing the update, so that our data is as much as possible in sync with the retailers' daily updated data.

Then we shall run the two weeks test period in which slight changes might need to be made and in the periods of daily update the coder shall be available on standby in case anything goes wrong.


If you have the required skills and experience, please bid on the project.
Related categories: Python Web Scraping Data Mining