Advanced Web Crawler and Scraper Development
Budget: $250 – $750 CAD
Hello
I need a web crawler/scraper which operates on the darkweb and clearweb,
- ignores dontfollow
-can use a configurable number of concurrent threads, indexes everything it finds into a backend database (including the content of websites, archives, text, documents, etc)
-can be set to seek out and crawl sites based on certain keywords and use different search engines to find sites to search
-can generate random darkwebaddresses to find and crawl sites
-has cli and gui interface, supports commands being sent over ssh or has a server and client version?
-supports configuring crawl depth, concurrent crawling threads, and other throttling settings
-supports whitelisting and black listing domains, tlds
-ideally has a basic gui but I'd be happy to help design that and its not 100% nessesary so long as there is an effective, interactive cli interface
-builds a searchable index of domains/subdomains/IPs/other basic info about the sites and an index of content/information within sites crawled
-configurable options on what data to index by cli flags and by regex
-configurable options on what data to index by cli flags and by regex
-api integration woiuld be great, allowing us to punch in our shodan,censys, criminalIP, etc api keys to use those sites as data sources but also to be able to use api to reeetrieve data from major OSINT sources such as facebook would be excellent but is not strictly necessary at this point
-being able to operate as a distributed task
-ability to connect multiple different cloud and local storage devices for indexed information
-a web interface for the GUI which features index information & stats/ configiuration options and the ability to search the indexed information for keywords or filter the information by type, file extension, regex, domain
-we will be wanting to integrate data analytics and ML/ai into the program (possibly down the road) and would like the program to get better at sniffing out sensitive data as it continueously crawls, scrapes and indexes its way through the darkweb.
we are running a darkweb monitoring service which we want to expand the capabilities of to improve the service offered to our clients so basically what we need the software to do is to find and database any sensitive information/personal information/financial information/business information/ identity documents/related material which has been leaked onto the darkweb (or clearweb) in order to notify our clients of their information being vulnerable. due to the nature of the service we provide we need this solution to be extremely through at finding this information but also extremely secure with indexing and storing it
I need a web crawler/scraper which operates on the darkweb and clearweb,
- ignores dontfollow
-can use a configurable number of concurrent threads, indexes everything it finds into a backend database (including the content of websites, archives, text, documents, etc)
-can be set to seek out and crawl sites based on certain keywords and use different search engines to find sites to search
-can generate random darkwebaddresses to find and crawl sites
-has cli and gui interface, supports commands being sent over ssh or has a server and client version?
-supports configuring crawl depth, concurrent crawling threads, and other throttling settings
-supports whitelisting and black listing domains, tlds
-ideally has a basic gui but I'd be happy to help design that and its not 100% nessesary so long as there is an effective, interactive cli interface
-builds a searchable index of domains/subdomains/IPs/other basic info about the sites and an index of content/information within sites crawled
-configurable options on what data to index by cli flags and by regex
-configurable options on what data to index by cli flags and by regex
-api integration woiuld be great, allowing us to punch in our shodan,censys, criminalIP, etc api keys to use those sites as data sources but also to be able to use api to reeetrieve data from major OSINT sources such as facebook would be excellent but is not strictly necessary at this point
-being able to operate as a distributed task
-ability to connect multiple different cloud and local storage devices for indexed information
-a web interface for the GUI which features index information & stats/ configiuration options and the ability to search the indexed information for keywords or filter the information by type, file extension, regex, domain
-we will be wanting to integrate data analytics and ML/ai into the program (possibly down the road) and would like the program to get better at sniffing out sensitive data as it continueously crawls, scrapes and indexes its way through the darkweb.
we are running a darkweb monitoring service which we want to expand the capabilities of to improve the service offered to our clients so basically what we need the software to do is to find and database any sensitive information/personal information/financial information/business information/ identity documents/related material which has been leaked onto the darkweb (or clearweb) in order to notify our clients of their information being vulnerable. due to the nature of the service we provide we need this solution to be extremely through at finding this information but also extremely secure with indexing and storing it