MOSHI-SPEED: ML/NLP Engineer — Voice-to-Voice LLM Fine-Tuning (Moshi/PersonaPlex) | Datasets Ready

Job ID: 40571903

Budget: $750 – $1,000 USD

Start your proposal with the word "MOSHI-SPEED". If you don't, your bid will be automatically deleted as a bot response.
We are looking for a Senior ML/AI Engineer to fine-tune our voice assistant built on the Moshi / PersonaPlex streaming audio architecture. The goal is to extend the model to support Russian + Uzbek (Latin/Cyrillic) while maintaining zero-latency native switching.
What is already done:
All datasets (instruction/dialogue formats) are fully prepared, cleaned, and ready for training. No data collection or labeling is required.
Scope of Work:
Tokenizer Audit & Extension: Optimize the tokenizer for the Uzbek language (both scripts).
SFT Fine-Tuning: Apply LoRA/QLoRA adjustments using the provided datasets. Crucially, updates must change weights without adding new layers.
Persona Alignment & Latency Control: Ensure character and tone consistency across RU/UZ and strictly maintain the current sub-120ms total latency (10.6ms inference step).
Deliverables:
Fine-tuned adapters (LoRA/QLoRA) or a fully merged model.
Clean, documented training scripts and evaluation reports.