Scrape Contact Information Extraction from 250M Websites using AI

Job ID: 34317537

Budget: $100 – $700 USD

I'm looking to crawl 250 Million unique URLs and scrape the contact information from them using AI / NLP / Regex. The scrape should check the seed URL, about page, contact page, support, and faq for the information

Below is the information I'm looking to scrape. All the information is unstructured and in natural language.

Business / Brand name
Business / Brand description
Is business (true or false)
Website Category (ie e-commerce, blog,etc)
Website subcategory( ie clothing store)
Emails
Address (using NLP *not just the <address> tag)
Phone Numbers
Twitter
Youtube
Facebook
LinkedIn
Instagram

I would like the export to be exported into a CSV/database.
you will have to use a combination of regex and NLP to extract the information

Please only apply if you have experience with AI as that will be needed for the job.