Scrape Contact Information Extraction from 250M Websites using AI
Budget: $100 – $700 USD
I'm looking to crawl 250 Million unique URLs and scrape the contact information from them using AI / NLP / Regex. The scrape should check the seed URL, about page, contact page, support, and faq for the information
Below is the information I'm looking to scrape. All the information is unstructured and in natural language.
Business / Brand name
Business / Brand description
Is business (true or false)
Website Category (ie e-commerce, blog,etc)
Website subcategory( ie clothing store)
Emails
Address (using NLP *not just the <address> tag)
Phone Numbers
Twitter
Youtube
Facebook
LinkedIn
Instagram
I would like the export to be exported into a CSV/database.
you will have to use a combination of regex and NLP to extract the information
Please only apply if you have experience with AI as that will be needed for the job.
Below is the information I'm looking to scrape. All the information is unstructured and in natural language.
Business / Brand name
Business / Brand description
Is business (true or false)
Website Category (ie e-commerce, blog,etc)
Website subcategory( ie clothing store)
Emails
Address (using NLP *not just the <address> tag)
Phone Numbers
Youtube
I would like the export to be exported into a CSV/database.
you will have to use a combination of regex and NLP to extract the information
Please only apply if you have experience with AI as that will be needed for the job.