Development of a Web Data Extraction Software for Company Information
Budget: $15 – $25 USD
Im looking to hire an experienced software developer (or a small team) to create a custom software solution that can automatically extract comprehensive company data from publicly available sources on the internet.
The software should be capable of retrieving and structuring key information about companies, including but not limited to:
Company Name
Company Address (Street, City, ZIP Code, Country, )
Phone Number(s) (landline and mobile)
Email Address(es)
Website URL
CEO / Managing Director Name
Key Employees (with titles/roles)
Social Media Links (e.g., LinkedIn, Facebook)
Industry Sector / Business Type
Company Registration Number (if available)
Additional public data (such as number of employees, founding year, revenue estimates, etc.)
Functional Requirements:
A user-friendly interface or dashboard to input search parameters (e.g., company name, country, industry).
The ability to scrape and extract data from various public sources, including:
Company directories
LinkedIn
Google Maps
Chamber of commerce databases
Company websites
Data should be automatically structured and stored in a database or exportable format (CSV, Excel, etc.).
The software should include deduplication logic to avoid duplicate records.
Optionally, ability to run bulk queries (e.g., search for all construction companies in a specific city).
Compliance with legal and ethical web scraping practices (e.g., rate limiting, robots.txt respect where necessary).
Technical Requirements:
Can be built as a desktop or web-based application.
Preferred programming languages: Python, JavaScript (Node.js), or similar.
Use of robust libraries or frameworks (e.g., BeautifulSoup, Scrapy, Puppeteer, Selenium).
Scalable architecture to add more data sources in the future.
Integration with APIs where possible (e.g., LinkedIn API, Google Places API).
Ideal Candidate Profile:
Proven experience in web scraping and data extraction projects.
Familiar with data parsing, handling anti-bot mechanisms (e.g., CAPTCHAs, proxies).
Strong attention to detail and clean code practices.
Ability to suggest and implement the best tools/tech stack for this project.
Optional: Experience in machine learning or NLP for data cleaning and enrichment.
Project Goal:
The end goal is to have a reliable tool that can help automate the process of gathering relevant business information at scale for research, business development, and outreach purposes.
The software should be capable of retrieving and structuring key information about companies, including but not limited to:
Company Name
Company Address (Street, City, ZIP Code, Country, )
Phone Number(s) (landline and mobile)
Email Address(es)
Website URL
CEO / Managing Director Name
Key Employees (with titles/roles)
Social Media Links (e.g., LinkedIn, Facebook)
Industry Sector / Business Type
Company Registration Number (if available)
Additional public data (such as number of employees, founding year, revenue estimates, etc.)
Functional Requirements:
A user-friendly interface or dashboard to input search parameters (e.g., company name, country, industry).
The ability to scrape and extract data from various public sources, including:
Company directories
Google Maps
Chamber of commerce databases
Company websites
Data should be automatically structured and stored in a database or exportable format (CSV, Excel, etc.).
The software should include deduplication logic to avoid duplicate records.
Optionally, ability to run bulk queries (e.g., search for all construction companies in a specific city).
Compliance with legal and ethical web scraping practices (e.g., rate limiting, robots.txt respect where necessary).
Technical Requirements:
Can be built as a desktop or web-based application.
Preferred programming languages: Python, JavaScript (Node.js), or similar.
Use of robust libraries or frameworks (e.g., BeautifulSoup, Scrapy, Puppeteer, Selenium).
Scalable architecture to add more data sources in the future.
Integration with APIs where possible (e.g., LinkedIn API, Google Places API).
Ideal Candidate Profile:
Proven experience in web scraping and data extraction projects.
Familiar with data parsing, handling anti-bot mechanisms (e.g., CAPTCHAs, proxies).
Strong attention to detail and clean code practices.
Ability to suggest and implement the best tools/tech stack for this project.
Optional: Experience in machine learning or NLP for data cleaning and enrichment.
Project Goal:
The end goal is to have a reliable tool that can help automate the process of gathering relevant business information at scale for research, business development, and outreach purposes.