Design and Implement RAG System Locally

Job ID: 40028270

Budget: $15 – $25 USD

On-Premise RAG System
1. Introduction and Project Scope
We are seeking proposals from full-stack technology provider to design, implement, and transfer a Confidential Retrieval-Augmented Generation (RAG) System based on a proprietary knowledge corpus. The primary goal is to enable high-performance semantic search and, optionally, question answering over our internal documentation, prioritizing data security and system flexibility.

Requirement:
• Knowledge Corpus Size: Approximately 1,000 documents (low growth rate).
• Document Formats: Diverse, including PDF (native and scanned), DOCX, PPTX, and XLS.
• Languages: English and Spanish (Multilingual capability is mandatory).
• Peak Concurrency: The system must reliably support a minimum of 10 concurrent users with an end-to-end response latency of under 5 seconds.

2. Mandatory Technical Requirements
The proposed solution must adhere strictly to the following criteria:
• 100% On-Premise Deployment: No reliance on public cloud services (AWS, Azure, GCP) for model inference, data storage, or vector databases. Data must never leave our local network.
• Deployment Architecture: The solution must be delivered using a containerization approach (e.g., Docker, Kubernetes, or similar) to ensure simplified, reproducible deployment and maintainability.
• High Performance via Pre-processing: The architecture must prioritize performance by heavily leveraging an offline pre-processing pipeline (document parsing, chunking, topic modeling, vector generation) to minimize latency during live queries.
• Access Control and Security: The system must integrate with our existing access management infrastructure (to be specified later, e.g., LDAP/OAuth) to ensure that users can only retrieve information from documents to which they have explicit read permissions.

3. Supplier Proposal Request (Technical & Stack)
We request that the supplier propose their preferred full-stack solution, detailing the key components and rationale. We are seeking a flexible, scalable, and modern architecture.
3.1. Proposed Technology Stack
Please specify your proposed stack for the following components. While we encourage flexibility, reference examples include LangChain, LlamaIndex, RAGFlow, ,Minima, FAISS, ChromaDB, Llama 3, Mistral, etc.

• Proposed Technology / Model
• Justification (Performance, Multilinguality, etc.)
• Document Parsing & Chunking
• Multilingual Embedding Model
• Vector Database (DBV)
• Topic Modeling / Clustering Method
• LLM (If used for generation)
• Inference Server (for concurrency)
3.2. Architecture and Data Flow
Please provide a high-level diagram illustrating the proposed architecture, detailing the flow from document ingestion to the user query response, with explicit attention to:
• The pre-processing phase (vector creation and topic modeling).
• The access control mechanism applied during the retrieval phase.
• How concurrency will be managed on the specified hardware.

4. System Flexibility, Scalability, and Transfer
4.1. Flexibility and Scalability
1. How will the solution allow us to easily swap or update the LLM and/or the embedding model in the future without requiring a major re-architecture?
2. How is the system designed to scale horizontally should the document corpus grow significantly (e.g., from 1,000 to 10,000 documents) or if the concurrency requirement increases?
4.2. Knowledge Transfer and Ownership
The final and non-negotiable condition is the complete transfer of the solution and all intellectual property (IP) to our company.
1. Please detail your knowledge transfer plan. This must include training sessions, documentation, and handover of all source code, configuration files, and deployment scripts (Dockerfiles, Kubernetes YAMLs, etc.).
2. Provide confirmation that all project IP, excluding any established commercial licensing of your company (if applicable), will belong entirely to our company upon completion.

5. Commercial and Pricing
Please provide a detailed breakdown of costs. We can divide the project into phases and provide a rough estimate in hours for each one, for example:

6. Supplier Information, relevant experience.
Related categories: AI Model Development AI Research