Evaluate RAG PDF Search Solution
Budget: $30 – $250 USD
Project Overview
I am looking for assistance evaluating and selecting a solution for semantic search across approximately 100,000 PDF documents.
The system should support a retrieval-based architecture using embeddings and a vector database. Optionally, it may support a RAG-style workflow with an LLM, but the LLM must be optional and able to be disabled.
When the LLM is disabled, the system should return relevant document excerpts only, avoiding generated answers.
The main goal of the system is reliable document search, not a conversational chatbot.
Requirements
The proposed solution should:
• Preferably use open-source or free-to-use tools with no mandatory licensing costs.
• Have moderate computational requirements, suitable for running on a workstation or small server.
• Use embeddings and a vector database for semantic search.
• Include a GUI suitable for non-technical users.
• Allow users to search documents and view relevant excerpts.
• Clearly indicate document source and page number in the results.
• Optionally allow LLM-generated answers to be enabled or disabled.
Document Ingestion
Documents will not be uploaded by users. PDFs are produced internally by the organization and will be indexed through a separate batch ingestion process (text extraction, chunking and embeddings generation).
The system should therefore support:
• batch indexing of documents
• updating the index when new PDFs are added
End users will only search the indexed document corpus.
Expected Deliverables
1.A comparison table of possible tools (pros/cons, complexity, scalability).
2.Installation instructions for the selected solution.
3.A short user manual explaining:
◦ how documents are indexed
◦ how users perform searches
◦ how to enable/disable LLM responses.
4.Brief support during initial setup.
Preferred Experience
Experience with:
•semantic search systems
•vector databases
•embeddings-based retrieval
•open-source tools.
Proposal
Please include:
•which tools you would recommend for this project
•a brief example of a similar system you have worked on.
To confirm that you have read the project description, please start your proposal with the word VECTOR and mention one vector database you have used before
I am looking for assistance evaluating and selecting a solution for semantic search across approximately 100,000 PDF documents.
The system should support a retrieval-based architecture using embeddings and a vector database. Optionally, it may support a RAG-style workflow with an LLM, but the LLM must be optional and able to be disabled.
When the LLM is disabled, the system should return relevant document excerpts only, avoiding generated answers.
The main goal of the system is reliable document search, not a conversational chatbot.
Requirements
The proposed solution should:
• Preferably use open-source or free-to-use tools with no mandatory licensing costs.
• Have moderate computational requirements, suitable for running on a workstation or small server.
• Use embeddings and a vector database for semantic search.
• Include a GUI suitable for non-technical users.
• Allow users to search documents and view relevant excerpts.
• Clearly indicate document source and page number in the results.
• Optionally allow LLM-generated answers to be enabled or disabled.
Document Ingestion
Documents will not be uploaded by users. PDFs are produced internally by the organization and will be indexed through a separate batch ingestion process (text extraction, chunking and embeddings generation).
The system should therefore support:
• batch indexing of documents
• updating the index when new PDFs are added
End users will only search the indexed document corpus.
Expected Deliverables
1.A comparison table of possible tools (pros/cons, complexity, scalability).
2.Installation instructions for the selected solution.
3.A short user manual explaining:
◦ how documents are indexed
◦ how users perform searches
◦ how to enable/disable LLM responses.
4.Brief support during initial setup.
Preferred Experience
Experience with:
•semantic search systems
•vector databases
•embeddings-based retrieval
•open-source tools.
Proposal
Please include:
•which tools you would recommend for this project
•a brief example of a similar system you have worked on.
To confirm that you have read the project description, please start your proposal with the word VECTOR and mention one vector database you have used before