Email Extractor (Python, AI, Web scrap)

Job ID: 37182028

Budget: ₹37,500 – ₹75,000 INR

I am looking for a skilled Python developer with experience in AI and web scraping to create an email extractor tool for research and data analysis purposes. The ideal candidate will have a strong understanding of data extraction techniques and be able to develop a system that can efficiently extract emails from websites.

Requirements:
- Proficiency in Python programming language
- Experience in web scraping and data extraction
- Knowledge of AI techniques for data analysis
- Ability to separate emails based on specific industry criteria
- Familiarity with handling large volumes of data (40,000+ emails)

Responsibilities:
- Develop a Python-based email extractor tool
- Implement web scraping techniques to extract emails from websites
- Modify the system to separate emails based on industry criteria
- Ensure the system can handle large volumes of data efficiently

If you have the necessary skills and experience, please submit your proposal. This is a great opportunity to work on a project that involves AI, web scraping, and data analysis.

Project Details:-
E-mail ID extraction success rate should be 100% where email IDs are present on the website.


If a website has more than one Email ID then the software should collect all email IDs from the website with a maximum count of two email IDs. The data collection should be excluding the pages like Blog, News, Article and product pages.

Along with collecting email IDs, the software should be able to collect Facebook business page URLs from the same website. It will also collect facebook urls even if no email IDs are found on the website

If no email IDs or Facebook business page URLs are found then the software should report Not Found.

Extraction software should have start, stop and pause option.

Live view of data extraction should be displayed on the software home screen with auto scrolling

Export option should be always available whenever we wish to stop and collect the data.

If the internet connection is interrupted then it should auto pause the collection and make the data available for download whatever it has extracted and it should allow us to resume the collection from the same row once internet reconnected.

There should be a progress bar to show the progress on how far the data’s have been collected from the full count of data. If we have uploaded a sheet of 5000 data and the collection count is at 2456 then it will show 2456 out of 5000 with % of success rate and failure rate.

Approximately a sheet of 10000-12000 websites to be extracted in one hour time with an internet bandwidth speed of 100 Mbps. The time taken for a full sheet collection may get an exception of 25% extra time depending upon the data.

A data sheet of 40000-50000 website should take approximate 4 hours of time with an exception of 25% extra time.