Business Contact Data Scraping

Job ID: 40505481

Budget: $750 – $1,500 USD

I need to build a clean, comprehensive dataset of business contact details pulled from multiple online directories as well as the relevant SOS (Secretary-of-State) websites. The scale is large enough that you should already be comfortable managing pagination, dynamic content, hidden API calls, and the occasional login or captcha.

Here is what the engagement will involve:

• Source coverage – start with the directories and SOS sites I specify; be prepared to expand if a source proves thin or misses key fields.
• Field accuracy – at minimum I expect company name, full postal address, phone, email, and any registration numbers available on the SOS records.
• Data hygiene – deduplicate across sources, normalise addresses, and flag obviously invalid phones or emails.
• Delivery – a single, well-structured CSV file ready for immediate import. Alongside the file, include the scripts or notebooks you used (Python + BeautifulSoup/Scrapy/Selenium or similar) and a short README so I can rerun the pipeline later.

Acceptance criteria
1. At least 95 % of rows must contain all mandatory fields.
2. No more than 2 % duplicate businesses when matching on name + address.
3. Scripts re-create the same CSV on a fresh machine with only the listed dependencies.

If this matches your expertise in large-scale scraping, authenticated workflows and data cleaning, let’s discuss timelines and any edge cases you foresee.