Scalable Web Crawler & Backend for B2B Search Engine

Job ID: 39494218

Budget: ₹37,500 – ₹75,000 INR

Project Title: Build a Scalable Web Crawler + Backend System for B2B Database Search Engine (Apollo.io Style)


---

Project Overview:

We are looking for an experienced Full-Stack Developer or Data Engineer/Web Crawler Expert to build the backend for a B2B database search engine — similar in concept to Apollo.io or ZoomInfo — that focuses on collecting verified contact data (Email, Phone, LinkedIn, etc.) for CEOs, Founders, and Decision Makers across industries and geographies.


---

Project Scope:

Phase 1 Goal:

Collect and organize 1 million verified B2B contacts through intelligent web crawling/scraping.

Focus on scraping decision-makers only (CEOs, Founders, Directors).

Extract important data points:

Name

Job Title

Company

Industry

Business Email (high priority)

Phone number (if public)

Location

LinkedIn / Website URLs




---

Functional Requirements:

Web Crawler Requirements:

Crawl professional and public sources like:

LinkedIn (indirect, via company page / Google dorks / profiles)

Company websites

Business directories

Google Maps listings

Crunchbase, AngelList, etc.


Should auto-detect duplicate records.

Should use rotating proxies / IP management for high success rate.

Must bypass captchas and block mechanisms intelligently.

Provide scraping logs and success rate analytics.

Database & Search Engine:

Save data in structured format (PostgreSQL or MongoDB).

Full-text search engine for users to search by:

Industry

Title

Location

Keyword


Build a backend API for future frontend access.


Backend Features:

Superadmin dashboard to view stats, data growth, crawling health.

Secure user authentication & roles (admin, analyst, etc.)

Scalable architecture for future frontend or SaaS model.

Data export options: Excel / CSV / JSON.



---

Tech Stack Preference (Flexible):

Backend: Python (FastAPI / Django) or Node.js

Crawler: Scrapy / Playwright / Selenium with proxy rotation

Database: PostgreSQL or MongoDB

Search Engine: ElasticSearch (for filtering and speed)

Infrastructure: AWS / DigitalOcean / GCP

Authentication: JWT or OAuth



---

Accuracy Requirement:

Minimum 90% accuracy on email formatting and job roles.

If possible, use email verification API integration (like NeverBounce, Hunter.io, or Zerobounce).



---

Timeline:

Phase 1: Setup + Crawl 1 million leads in 6-8 weeks max

Will continue project in later phases if successful (frontend, SaaS model, client login)