Development of a Web Data Extraction Software for Company Information

Job ID: 39256550

Budget: $15 – $25 USD

Im looking to hire an experienced software developer (or a small team) to create a custom software solution that can automatically extract comprehensive company data from publicly available sources on the internet.

The software should be capable of retrieving and structuring key information about companies, including but not limited to:

Company Name

Company Address (Street, City, ZIP Code, Country, )

Phone Number(s) (landline and mobile)

Email Address(es)

Website URL

CEO / Managing Director Name

Key Employees (with titles/roles)

Social Media Links (e.g., LinkedIn, Facebook)

Industry Sector / Business Type

Company Registration Number (if available)

Additional public data (such as number of employees, founding year, revenue estimates, etc.)

Functional Requirements:
A user-friendly interface or dashboard to input search parameters (e.g., company name, country, industry).

The ability to scrape and extract data from various public sources, including:

Company directories

LinkedIn

Google Maps

Chamber of commerce databases

Company websites

Data should be automatically structured and stored in a database or exportable format (CSV, Excel, etc.).

The software should include deduplication logic to avoid duplicate records.

Optionally, ability to run bulk queries (e.g., search for all construction companies in a specific city).

Compliance with legal and ethical web scraping practices (e.g., rate limiting, robots.txt respect where necessary).

Technical Requirements:
Can be built as a desktop or web-based application.

Preferred programming languages: Python, JavaScript (Node.js), or similar.

Use of robust libraries or frameworks (e.g., BeautifulSoup, Scrapy, Puppeteer, Selenium).

Scalable architecture to add more data sources in the future.

Integration with APIs where possible (e.g., LinkedIn API, Google Places API).

Ideal Candidate Profile:
Proven experience in web scraping and data extraction projects.

Familiar with data parsing, handling anti-bot mechanisms (e.g., CAPTCHAs, proxies).

Strong attention to detail and clean code practices.

Ability to suggest and implement the best tools/tech stack for this project.

Optional: Experience in machine learning or NLP for data cleaning and enrichment.

Project Goal:
The end goal is to have a reliable tool that can help automate the process of gathering relevant business information at scale for research, business development, and outreach purposes.