Advanced Web Scraping Specialist Needed for Magento Website with Machine Learning and BeautifulSoup
Budget: $6 – $12 USD
We are currently facing challenges in scraping specific components from our Magento website. Our existing setup utilizes BeautifulSoup and Selenium, but we're encountering difficulties in accurately identifying the target components, such as image cards, tables, etc., based on predefined criteria. The main issue stems from the complexity of the website's structure, where the code primarily consists of divs, and the JavaScript and CSS are heavily compressed.
This complexity has led to a situation where adjustments made to the scraping logic work on some pages but fail or scrape incorrect data on others. With a vast website encompassing thousands of pages, manual adjustments are not feasible and often lead to a breakdown in other parts of the site.
We are looking for a skilled freelancer who can implement advanced techniques, potentially involving machine learning, to enhance our web scraping capabilities. The primary task will be to develop a system that can intelligently identify or tag complex web components that are currently challenging to scrape. This system should adapt to the diverse structures of different pages and reliably extract the required data.
Key Responsibilities:
Analyze the current web scraping setup and identify the key challenges in extracting specific components.
Develop and integrate a solution, possibly using machine learning, to accurately identify and tag complex web components.
Ensure that the solution is scalable and can handle thousands of pages with varying structures.
Conduct thorough testing to ensure accuracy and efficiency in data extraction across all pages.
Collaborate with our team to integrate the solution into the existing web scraping framework.
Required Skills and Qualifications:
Proven experience in web scraping, particularly with BeautifulSoup and Selenium.
Strong background in machine learning, with the ability to apply ML techniques to web scraping.
Familiarity with Magento websites and their structure.
Excellent problem-solving skills and attention to detail.
Ability to work independently and deliver effective solutions within deadlines.
This complexity has led to a situation where adjustments made to the scraping logic work on some pages but fail or scrape incorrect data on others. With a vast website encompassing thousands of pages, manual adjustments are not feasible and often lead to a breakdown in other parts of the site.
We are looking for a skilled freelancer who can implement advanced techniques, potentially involving machine learning, to enhance our web scraping capabilities. The primary task will be to develop a system that can intelligently identify or tag complex web components that are currently challenging to scrape. This system should adapt to the diverse structures of different pages and reliably extract the required data.
Key Responsibilities:
Analyze the current web scraping setup and identify the key challenges in extracting specific components.
Develop and integrate a solution, possibly using machine learning, to accurately identify and tag complex web components.
Ensure that the solution is scalable and can handle thousands of pages with varying structures.
Conduct thorough testing to ensure accuracy and efficiency in data extraction across all pages.
Collaborate with our team to integrate the solution into the existing web scraping framework.
Required Skills and Qualifications:
Proven experience in web scraping, particularly with BeautifulSoup and Selenium.
Strong background in machine learning, with the ability to apply ML techniques to web scraping.
Familiarity with Magento websites and their structure.
Excellent problem-solving skills and attention to detail.
Ability to work independently and deliver effective solutions within deadlines.