Scrape 900K Leads Dataset

Job ID: 39776948

Budget: $3,000 – $5,000 USD

need an experienced scraper who can pull roughly 900,000 records from two separate sites where the data appears through a mix of JavaScript-rendered pages and direct API calls. The goal is to capture clean, ready-to-use lead information—name, phone, email, location and any other contact-related fields the sites expose—without duplicates or corrupted rows.

Because both sources are heavily scripted and protected, you will have to rely on headless browsers or playwright/puppeteer-style tooling, smart proxy rotation, and reliable captcha-solving. I am comfortable with Python (Scrapy, Selenium, BeautifulSoup) or Node.js (Puppeteer, Playwright); choose whichever stack lets you move fastest while staying stable under load.

Deliverables
• A small verification sample of about 1,000 lines so I can check field coverage and formatting
• The complete, de-duplicated dataset (~900K rows) in the structured format of your choice—CSV or Excel both work for me, just keep each column clearly labelled

I am open on budget and timeline, but please outline:
1. The approach you plan to take against anti-bot measures
2. How long the run will take end-to-end once the sample is approved
3. Any checkpoints you suggest to guarantee quality before final delivery

If you have handled similarly sized jobs before, sharing a brief reference or example will help me award quickly.