Enterprise RAG Chatbot Development
Budget: $750 – $1,500 USD
I need a skilled developer to build and optimize a Retrieval-Augmented Generation (RAG) chatbot for our enterprise environment. The chatbot will be primarily used for customer support and internal team assistance, and needs to be integrated with our website.
Key Requirements:
- High accuracy in responses
- Fast response speed
- Scalability to handle
Ideal Skills and Experience:
- Expertise in AI and machine learning, particularly in chatbot development
- Experience with RAG models
- Strong background in web integration and scalability solutions
- Proven track record in delivering high-performance enterprise-level applications
1. Data collection & preprocessing : Load data from multiple formats (PDF, Word, PowerPoint, Excel, CSV, …) and set up Structure-aware parsing (headings, tables, lists), extract metadata from documents
2. Chunking & embedding techniques: Set up appropriate chunking techniques, SOTA. Or design chunking strategies suitable for each type of data. Use the best embedding models and store embeddings with versions for easy upgrades later
3. Integrate vector DB : Support metadata filtering, HNSW tuning, hybrid search (dense + BM25)
4. Optimize retrieval:
- Improve queries with HyDE, self-ask, step-back prompting
- Support multi-query, chain-of-thought retrieval
- Rerank results (cross-encoder, OpenAI rerank)
- Tune retrieval to achieve optimal precision and recall
5. Integrate LLM & build context
- Use the most efficient LLM API in terms of quality and cost. Flexible switch settings to LLM are cheaper if the query is simple.
- Can easily switch to self-hosted LLM
- Prompt design ensures accuracy and adherence to retrieved content
- Ability to display citation sources and assess reliability
- Support long contexts (sliding window, packing by token, summarization, …)
6. System evaluation & monitoring
- Reporting retrieval quality with benchmarks (recall, precision, nDCG)
- Assessing answer quality (accuracy, honesty, citation sources)
- Integrating dashboard/logs to monitor system quality
- Testing with real QA datasets
7. Infrastructure & scaling
- Backend using FastAPI / LangChain / LlamaIndex
- Support inference on GPU (Triton, OpenVINO if needed)
- API security (auth, rate limit, log)
- Cache effectively embedding and LLM results
8. Documentation & handover
- Clear codebase, detailed documentation and architecture diagram
- Internal engineer guidance system takeover
Desired roadmap:
- Phase 1 (4 weeks): have MVP
- Phase 2 (2 weeks): install advanced techniques to optimize performance, log & dashboard
- Phase 3: expand, handover, support maintenance if our customers have problems
Apply:
Please send us:
- RAG/chatbot related projects you have done (link/code if available)
- Proposed architecture and stack used
- Team structure
- Quotation (full package)
- Estimated implementation time
Key Requirements:
- High accuracy in responses
- Fast response speed
- Scalability to handle
Ideal Skills and Experience:
- Expertise in AI and machine learning, particularly in chatbot development
- Experience with RAG models
- Strong background in web integration and scalability solutions
- Proven track record in delivering high-performance enterprise-level applications
1. Data collection & preprocessing : Load data from multiple formats (PDF, Word, PowerPoint, Excel, CSV, …) and set up Structure-aware parsing (headings, tables, lists), extract metadata from documents
2. Chunking & embedding techniques: Set up appropriate chunking techniques, SOTA. Or design chunking strategies suitable for each type of data. Use the best embedding models and store embeddings with versions for easy upgrades later
3. Integrate vector DB : Support metadata filtering, HNSW tuning, hybrid search (dense + BM25)
4. Optimize retrieval:
- Improve queries with HyDE, self-ask, step-back prompting
- Support multi-query, chain-of-thought retrieval
- Rerank results (cross-encoder, OpenAI rerank)
- Tune retrieval to achieve optimal precision and recall
5. Integrate LLM & build context
- Use the most efficient LLM API in terms of quality and cost. Flexible switch settings to LLM are cheaper if the query is simple.
- Can easily switch to self-hosted LLM
- Prompt design ensures accuracy and adherence to retrieved content
- Ability to display citation sources and assess reliability
- Support long contexts (sliding window, packing by token, summarization, …)
6. System evaluation & monitoring
- Reporting retrieval quality with benchmarks (recall, precision, nDCG)
- Assessing answer quality (accuracy, honesty, citation sources)
- Integrating dashboard/logs to monitor system quality
- Testing with real QA datasets
7. Infrastructure & scaling
- Backend using FastAPI / LangChain / LlamaIndex
- Support inference on GPU (Triton, OpenVINO if needed)
- API security (auth, rate limit, log)
- Cache effectively embedding and LLM results
8. Documentation & handover
- Clear codebase, detailed documentation and architecture diagram
- Internal engineer guidance system takeover
Desired roadmap:
- Phase 1 (4 weeks): have MVP
- Phase 2 (2 weeks): install advanced techniques to optimize performance, log & dashboard
- Phase 3: expand, handover, support maintenance if our customers have problems
Apply:
Please send us:
- RAG/chatbot related projects you have done (link/code if available)
- Proposed architecture and stack used
- Team structure
- Quotation (full package)
- Estimated implementation time
Related categories:
Business, Accounting, Human Resources & Legal
Python
Excel
Web Scraping
Data Mining