Improve RAG Assistant Inference Speed

Job ID: 38457674

Budget: $30 – $250 USD

I'm in need of an expert to help me enhance the performance of my RAG assistant on my local PC. The assistant currently handles generative AI, specifically text generation, but I'm facing significant issues with its inference speed.

Key Tasks:
- Analyze the current setup and identify bottlenecks that are affecting the performance, especially the inference speed of the RAG assistant.
- Suggest and implement optimizations that will improve the inference speed without compromising the quality of the text generation.
- Test the enhancements thoroughly to ensure they are effective and reliable.

Ideal Skills and Experience:
- Proficiency in generative AI, particularly text generation, is a must.
- Extensive experience in optimizing inference speed for generative AI models.
- Strong background in performance tuning on local PCs.
- Excellent problem-solving and analytical skills.

I'm looking for a professional who can not only boost the inference speed of my RAG assistant but also provide insights into long-term strategies for maintaining optimal performance.
Related categories: LLaMA Generative AI Streamlit