Speech-to-Speech RAG Application
Budget: $30 – $250 USD
I'm looking for a Python LLM developer who can create a speech-to-speech RAG application for personal use.
The project is basically mimicking the opensource git project:
Medium: https://medium.com/@rossashman/meet-emma-voice-to-voice-rag-84dab1a5ef1b
Code: https://github.com/ploncker/voice_to_voice_rag/blob/main/gradio_voice2voice_rag.py
Ideal skills for this job:
- Generative AI use (especially proficiency in use of LLM eg. Llama-3.1-8b for RAG application)
- Proficient in compression of llm compressor and use of vLLM for inference. No OpenAI or any paid model.
- Proficient in desktop and web-based application development
- Experience with speech recognition and speech synthesis technology
- Strong problem-solving skills and attention to detail.
Requirements in my opinion:
- An LLM Model for
- e.g Mistral-7B , LlaMA-3.1. compression
- Speech-to-Speech Application
- Proficient in the use of open-source speech application: e.g Whisper
- Nice Flask/Django interface
The application is an opensource project and is simply. Code already available. Your challenge would be using Big LLM model. Compression is not necessary if you can get it to work with Ollama - local Llama.
The project is basically mimicking the opensource git project:
Medium: https://medium.com/@rossashman/meet-emma-voice-to-voice-rag-84dab1a5ef1b
Code: https://github.com/ploncker/voice_to_voice_rag/blob/main/gradio_voice2voice_rag.py
Ideal skills for this job:
- Generative AI use (especially proficiency in use of LLM eg. Llama-3.1-8b for RAG application)
- Proficient in compression of llm compressor and use of vLLM for inference. No OpenAI or any paid model.
- Proficient in desktop and web-based application development
- Experience with speech recognition and speech synthesis technology
- Strong problem-solving skills and attention to detail.
Requirements in my opinion:
- An LLM Model for
- e.g Mistral-7B , LlaMA-3.1. compression
- Speech-to-Speech Application
- Proficient in the use of open-source speech application: e.g Whisper
- Nice Flask/Django interface
The application is an opensource project and is simply. Code already available. Your challenge would be using Big LLM model. Compression is not necessary if you can get it to work with Ollama - local Llama.
Related categories:
Python
Machine Learning (ML)
Flask
Large Language Models (LLMs)
Retrieval-Augemented Generation