MLOps Engineer for Deployment of LLMs

Job ID: 38601957

Budget: $3,000 – $5,000 USD

MLOps Engineers welcome, we are looking to deploy in 3-4 weeks:

Text-to-Text Chat:

- Deploy Chat LLM

- Currently evaluating Llama V2 Uncensored (TheBloke) or Mistral as potential options

- Implement Long Term Memory Solution to retain conversational context over time, this could be achieved through knowledge graphs or vector databases to enhance chat experience

- The chat system involves keyword matching to determine if user is asking for a picture "show me" or "take a picture"

- Prepare detailed workflow integration steps with our backend developers keeping context above in mind

Text-to-Image Workflow:

- Face continuity workflow deployment, deploy the ADetailer face continuity workflow to ensure high quality facial consistency for image outputs (we already have prompt framework to generate faces and bodies) Stop here, this is a test, to know you have read our description please put "HONEYBADGER" at the top of your proposal.

- Deploy RealVisXL V4.0 BAE, host on a reliable cloud platform that supports external API access for inference, with robust monitoring to ensure high availability and performance.

- This main functionality is used to create the initial preview of the AI person as well as a reaction to any chat initiated requests as detailed above in text to text milestone.

- Prepare detailed workflow integration steps with the backend developers keeping the context above in mind.

Good-to-Have Features

- Sentiment Analysis Integration: Implement sentiment analysis within the LLMs to refine response generation. This enhancement would add emotional intelligence to the chat model, improving user experience by dynamically adjusting responses based on sentiment.

NOTES:

We currently have a team of app developers, with existing services deployed on Netlify and Render.com. The ML services should be integrated via accessible inference endpoints, enabling seamless connection to our current backend APIs.