Custom Independent Search Engine Development

Job ID: 39031917

Budget: $200 – $500 USD

We are looking for an experienced developer or team to create a custom search engine from scratch. The goal is to build a fully functional search engine that operates independently, without relying on any third-party APIs (e.g., Google, DuckDuckGo, or Bing). In addition, the system will include a robust Web Crawler and Indexing System, along with an Admin Panel to manage the entire process.
This project will involve developing a crawler, setting up an indexing system, and building an admin panel to control, monitor, and optimize the search engine.
Key Requirements:
1. Web Crawler Development:
• Purpose: Develop a custom crawler to scrape and collect data from websites.
• Details: The crawler should follow robots.txt rules and efficiently crawl multiple websites, handling large-scale data.
• Features:
• Set crawl depth and frequency.
• Handle crawling for dynamic content (e.g., AJAX, JavaScript-heavy pages).
• Log crawl activities and errors.
2. Indexing System:
• Purpose: Develop an indexing system to organize the crawled data and make it searchable.
• Details: The system should efficiently store data and provide quick search results based on relevant parameters.
• Features:
• Manage and optimize the storage of crawled data (text, images, metadata).
• Re-index data when necessary.
• Clean-up and remove outdated data from the index.
• Develop custom ranking algorithms (similar to PageRank or keyword relevance).
3. Admin Panel:
• Purpose: Build an Admin Panel to allow easy management of the crawler and indexed data.
• Features:
• Crawler Settings: Control crawl frequency, depth, and site list.
• Index Management: View, edit, and delete indexed data.
• Crawl Logs: Display logs of all crawling activity, errors, and successful crawls.
• Re-indexing Options: Manually trigger re-indexing when data changes or needs to be refreshed.
• System Monitoring: View real-time stats about the crawling process and indexed data performance.
4. Search Engine Development:
• Purpose: Develop the frontend and backend of the search engine for users to search results from the indexed data.
• Features:
• Simple, user-friendly interface for searching indexed data.
• Fast and relevant search results.
• Results should be displayed with the most relevant information at the top (based on keyword match, ranking, etc.).
5. Database Integration:
• Purpose: Set up a database to store and manage crawled data and indexes.
• Features:
• Use a scalable database like MySQL, PostgreSQL, or MongoDB.
• Efficient query handling to fetch search results quickly.
6. Scalability & Performance:
• Goal: Ensure the entire system is optimized for performance, scalability, and reliability.
• Details: The search engine should handle a large volume of data and multiple simultaneous users.
Deliverables:
• Custom Web Crawler that scrapes and collects data from specified websites.
• Indexing System to organize and store crawled data efficiently.
• Admin Panel for managing crawler settings, viewing logs, and controlling the search engine.
• Search Engine (frontend and backend) that provides fast, relevant search results based on indexed data.
• Source code with comprehensive documentation.
• Deployment instructions and support.
Ideal Candidate:
• Proven experience in developing custom search engines or web crawlers.
• Strong knowledge of indexing systems and search algorithms.
• Expertise in backend development, preferably with PHP or Python.
• Familiarity with databases such as MySQL, PostgreSQL, or MongoDB.
• Experience in building scalable, high-performance systems.
• Previous work with admin panels and data management systems is a plus.

Additional Notes:
• Third-party APIs are not allowed for search results or data collection; everything must be built from scratch.
• The project should be completed within [timeframe], with regular progress updates.