NVIDIA PersonaPlex: Russian & Uzbek Optimization
Budget: $250 – $750 USD
Title: Fine-tuning NVIDIA PersonaPlex for Russian and Uzbek Languages
Project Overview:
We are developing a digital assistant based on the NVIDIA PersonaPlex reference architecture. Currently, the system is optimized for English. We need an expert to adapt and fine-tune the LLM core to support Russian and Uzbek languages using our proprietary dataset.
Technical Stack:
Framework: NVIDIA NeMo / NVIDIA Riva (PersonaPlex Pipeline).
Hardware: NVIDIA L40S GPU (48GB VRAM).
Optimization: TensorRT-LLM.
Scope of Work:
Baseline Assessment: Evaluate the current PersonaPlex LLM (e.g., Llama-3 or Mistral-based) for multilingual capabilities.
Tokenizer Adaptation: Analyze if the current tokenizer supports Uzbek (Latin/Cyrillic) and Russian. Perform vocabulary expansion if necessary to prevent high fragmentation.
Fine-tuning (SFT/LoRA): Perform Supervised Fine-Tuning (SFT) or LoRA/QLoRA using the provided dataset (Russian/Uzbek instructions and dialogues).
Optimization: Convert and optimize the fine-tuned model using TensorRT-LLM for low-latency inference on L40S.
Integration: Ensure the persona-based logic (system prompts, emotional cues) remains consistent in Russian and Uzbek.
Validation: Test the model for language switching, grammatical accuracy, and cultural relevance.
Requirements for Freelancer:
Proven experience with NVIDIA NeMo Framework and TensorRT-LLM.
Strong background in LLM Fine-tuning (SFT, PEFT).
Experience with multilingual models and tokenizer customization.
Access to appropriate compute or ability to work via remote access to our L40S environment.
Deliverables:
Fine-tuned weights (or LoRA adapters).
Optimized TensorRT-LLM engine.
Training scripts and documentation for future updates.
Project Overview:
We are developing a digital assistant based on the NVIDIA PersonaPlex reference architecture. Currently, the system is optimized for English. We need an expert to adapt and fine-tune the LLM core to support Russian and Uzbek languages using our proprietary dataset.
Technical Stack:
Framework: NVIDIA NeMo / NVIDIA Riva (PersonaPlex Pipeline).
Hardware: NVIDIA L40S GPU (48GB VRAM).
Optimization: TensorRT-LLM.
Scope of Work:
Baseline Assessment: Evaluate the current PersonaPlex LLM (e.g., Llama-3 or Mistral-based) for multilingual capabilities.
Tokenizer Adaptation: Analyze if the current tokenizer supports Uzbek (Latin/Cyrillic) and Russian. Perform vocabulary expansion if necessary to prevent high fragmentation.
Fine-tuning (SFT/LoRA): Perform Supervised Fine-Tuning (SFT) or LoRA/QLoRA using the provided dataset (Russian/Uzbek instructions and dialogues).
Optimization: Convert and optimize the fine-tuned model using TensorRT-LLM for low-latency inference on L40S.
Integration: Ensure the persona-based logic (system prompts, emotional cues) remains consistent in Russian and Uzbek.
Validation: Test the model for language switching, grammatical accuracy, and cultural relevance.
Requirements for Freelancer:
Proven experience with NVIDIA NeMo Framework and TensorRT-LLM.
Strong background in LLM Fine-tuning (SFT, PEFT).
Experience with multilingual models and tokenizer customization.
Access to appropriate compute or ability to work via remote access to our L40S environment.
Deliverables:
Fine-tuned weights (or LoRA adapters).
Optimized TensorRT-LLM engine.
Training scripts and documentation for future updates.
Related categories:
Translation
Russian Translator
Deep Learning
Natural Language Processing
AI Research
AI Development