Urgent Experienced Python Selenium Developer for Web Scraping Project
Budget: ₹1,500 – ₹12,500 INR
I am looking for a highly skilled and experienced Python developer with expertise in Selenium web scraping to assist with optimizing and enhancing our current web scraping project. This project involves collecting data from 3 specific websites for personal use, with a focus on improving speed and efficiency. The ideal candidate will have a strong background in web scraping, a deep understanding of Python and Selenium, and the ability to implement cost-effective solutions for running web scraping tasks.
Responsibilities:
Review and optimize existing Python Selenium web scraping scripts to improve speed and efficiency.
Implement headless browsing, and optimize browser settings to disable unnecessary elements such as images and JavaScript for faster page loads.
Introduce efficient data handling and storage mechanisms to manage the data collected from web scraping activities.
Explore and implement concurrency and parallel processing techniques to enable simultaneous scraping of multiple pages.
Advise on and implement caching strategies to reduce redundant downloads and speed up the scraping process.
Evaluate and recommend free or cost-effective cloud solutions for deploying the web scraping scripts, considering the potential for scaling and reducing latency.
Ensure the robustness of the web scraping solution, including error handling and avoiding IP bans or CAPTCHAs by target websites.
Document the code and provide guidance on running and maintaining the web scraping setup.
Requirements:
Proven experience with Python and Selenium for web scraping projects.
Strong understanding of web technologies (HTML, CSS, JavaScript) and how to interact with them using Selenium.
Experience with headless browsers and optimizing Selenium configurations for performance.
Knowledge of concurrency and parallel processing in Python (e.g., threading, asyncio).
Familiarity with caching strategies and their implementation in web scraping contexts.
Experience with cloud platforms (AWS, Google Cloud, etc.) and understanding of their free tiers and how to leverage them for web scraping tasks.
Ability to work independently, with minimal supervision, and deliver high-quality code.
Excellent problem-solving skills and attention to detail.
Strong communication skills, with fluency in English.
Nice to Have:
Experience with Docker and Selenium Grid for parallel processing.
Knowledge of other web scraping frameworks or libraries in Python such as Scrapy.
Responsibilities:
Review and optimize existing Python Selenium web scraping scripts to improve speed and efficiency.
Implement headless browsing, and optimize browser settings to disable unnecessary elements such as images and JavaScript for faster page loads.
Introduce efficient data handling and storage mechanisms to manage the data collected from web scraping activities.
Explore and implement concurrency and parallel processing techniques to enable simultaneous scraping of multiple pages.
Advise on and implement caching strategies to reduce redundant downloads and speed up the scraping process.
Evaluate and recommend free or cost-effective cloud solutions for deploying the web scraping scripts, considering the potential for scaling and reducing latency.
Ensure the robustness of the web scraping solution, including error handling and avoiding IP bans or CAPTCHAs by target websites.
Document the code and provide guidance on running and maintaining the web scraping setup.
Requirements:
Proven experience with Python and Selenium for web scraping projects.
Strong understanding of web technologies (HTML, CSS, JavaScript) and how to interact with them using Selenium.
Experience with headless browsers and optimizing Selenium configurations for performance.
Knowledge of concurrency and parallel processing in Python (e.g., threading, asyncio).
Familiarity with caching strategies and their implementation in web scraping contexts.
Experience with cloud platforms (AWS, Google Cloud, etc.) and understanding of their free tiers and how to leverage them for web scraping tasks.
Ability to work independently, with minimal supervision, and deliver high-quality code.
Excellent problem-solving skills and attention to detail.
Strong communication skills, with fluency in English.
Nice to Have:
Experience with Docker and Selenium Grid for parallel processing.
Knowledge of other web scraping frameworks or libraries in Python such as Scrapy.