Web Crawler Development for Site Indexing
Budget: $250 – $750 AUD
We are hiring a dev to build us a simple web crawler.
The desired tech stack we think is:
Features we like are similar to the Spider project (https://www.sphider.eu/about.php) which is free to download so perhaps we can start with this? -
Features:
Performs full text indexing.
Can index both static and dynamic pages.
Finds links in href, frame, area and meta tags, and can also follow links given in javascript as strings via window.location and window.open.
Respects robots.txt protocol, and nofollow and noindex tags.
Follows server side redirections.
Allows spidering to be limited by depth (ie maximum number of clicks from the starting page), by (sub)domain or by directory.
Allows spidering only the urls matching (or not matching) certain keywords or regular expressions.
Supports indexing of pdf and doc files (using external binaries for file conversion).
Allows resuming paused spidering.
Possbility to exclude common words from being indexed.
Also require a web based admin to manage with such features:
Includes a sophisticated web based administration interface
Supports indexing via a web interface as well as from commandline - easy to set up cron jobs.
Comprehensive site and search statistics
The desired tech stack we think is:
Features we like are similar to the Spider project (https://www.sphider.eu/about.php) which is free to download so perhaps we can start with this? -
Features:
Performs full text indexing.
Can index both static and dynamic pages.
Finds links in href, frame, area and meta tags, and can also follow links given in javascript as strings via window.location and window.open.
Respects robots.txt protocol, and nofollow and noindex tags.
Follows server side redirections.
Allows spidering to be limited by depth (ie maximum number of clicks from the starting page), by (sub)domain or by directory.
Allows spidering only the urls matching (or not matching) certain keywords or regular expressions.
Supports indexing of pdf and doc files (using external binaries for file conversion).
Allows resuming paused spidering.
Possbility to exclude common words from being indexed.
Also require a web based admin to manage with such features:
Includes a sophisticated web based administration interface
Supports indexing via a web interface as well as from commandline - easy to set up cron jobs.
Comprehensive site and search statistics