AI/Fullstack NLP / Speech Deep Learning Engineer (Core AI),AI Infrastructure / LLM DevOps Engineer -- 2
Budget: $750 – $1,500 USD
Lead AI / Fullstack Engineer — Project "AZIZA" (Voice-to-Voice AI)
Project Name: AZIZA
Format: Project-based / Remote (with access to local GPU clusters)
Tech Stack: PersonaPlex (Moshi-based architecture), PyTorch, TensorRT-LLM, FastAPI, WebRTC, Telegram Mini App (TMA).
Hardware Location: Uzbekistan & Turkey clusters powered by NVIDIA L40S
Project Overview
AZIZA is an innovative multimodal "Speech-to-Speech" (S2S) ecosystem designed to simulate natural human interaction. We are building an AI assistant that seamlessly transitions between roles: an expert tutor (Chemistry, History, Biology), an empathetic companion, and a simultaneous translator. By processing audio tokens directly, the system achieves unprecedented interaction speeds.
Current Status: The base model (English) is stable. We are now scaling to address regional specifics and deploying the solution within a high-tech application framework.
Key Responsibilities
1. Core AI & ML (Adaptation & Intelligence)
Multilingual Support: Lead cross-lingual fine-tuning to provide native-level support for Uzbek (including regional dialects), Kazakh, and Russian,Tadjik
Latency Optimization: Streamline inference pipelines to target a response latency of 180-300 milseconds.
Smart RAG (100 GB): Architect a vector knowledge base for educational materials, implementing a "triple-check" verification mechanism to eliminate hallucinations.
NVIDIA Stack: Optimize inference for L40S environments using vLLM, TensorRT-LLM, and INT4/FP8 quantization.
2. Telegram Mini App & Real-time Web
Audio Streaming: Implement low-latency real-time audio transmission via WebRTC / WebSockets (moving beyond standard voice message protocols).
Full-Duplex UI: Develop a frontend that supports interruptibility, allowing the AI to react instantly when the user speaks over it.
Billing: Integrate local payment gateways (Payme, Click) for subscription management.
3. Architecture & Infrastructure
Highload Design: Design a horizontally scalable system capable of handling high concurrent user loads.
Signal Processing: Implement software-based AEC (Acoustic Echo Cancellation) and noise suppression to ensure high-fidelity communication.
Traffic Localization: Optimize routing protocols to maximize performance within the TAS-IX network.
Candidate Requirements
AI / ML Engineering:
Proven experience with End-to-end (E2E) speech models (Moshi, AudioLM, or similar).
Deep proficiency in PyTorch and Transformer architectures.
Hands-on experience in Fine-tuning LLMs/S2S models for new language groups.
Expertise in CUDA 12.x and NVIDIA optimization libraries.
Fullstack Development:
Expert-level knowledge of WebRTC / WebSockets for real-time media streaming.
Demonstrated experience in developing Telegram Mini Apps (TMA).
Professional mastery of FastAPI and React / Next.js.
Strong understanding of the constraints and requirements of Low-latency systems.
Project Name: AZIZA
Format: Project-based / Remote (with access to local GPU clusters)
Tech Stack: PersonaPlex (Moshi-based architecture), PyTorch, TensorRT-LLM, FastAPI, WebRTC, Telegram Mini App (TMA).
Hardware Location: Uzbekistan & Turkey clusters powered by NVIDIA L40S
Project Overview
AZIZA is an innovative multimodal "Speech-to-Speech" (S2S) ecosystem designed to simulate natural human interaction. We are building an AI assistant that seamlessly transitions between roles: an expert tutor (Chemistry, History, Biology), an empathetic companion, and a simultaneous translator. By processing audio tokens directly, the system achieves unprecedented interaction speeds.
Current Status: The base model (English) is stable. We are now scaling to address regional specifics and deploying the solution within a high-tech application framework.
Key Responsibilities
1. Core AI & ML (Adaptation & Intelligence)
Multilingual Support: Lead cross-lingual fine-tuning to provide native-level support for Uzbek (including regional dialects), Kazakh, and Russian,Tadjik
Latency Optimization: Streamline inference pipelines to target a response latency of 180-300 milseconds.
Smart RAG (100 GB): Architect a vector knowledge base for educational materials, implementing a "triple-check" verification mechanism to eliminate hallucinations.
NVIDIA Stack: Optimize inference for L40S environments using vLLM, TensorRT-LLM, and INT4/FP8 quantization.
2. Telegram Mini App & Real-time Web
Audio Streaming: Implement low-latency real-time audio transmission via WebRTC / WebSockets (moving beyond standard voice message protocols).
Full-Duplex UI: Develop a frontend that supports interruptibility, allowing the AI to react instantly when the user speaks over it.
Billing: Integrate local payment gateways (Payme, Click) for subscription management.
3. Architecture & Infrastructure
Highload Design: Design a horizontally scalable system capable of handling high concurrent user loads.
Signal Processing: Implement software-based AEC (Acoustic Echo Cancellation) and noise suppression to ensure high-fidelity communication.
Traffic Localization: Optimize routing protocols to maximize performance within the TAS-IX network.
Candidate Requirements
AI / ML Engineering:
Proven experience with End-to-end (E2E) speech models (Moshi, AudioLM, or similar).
Deep proficiency in PyTorch and Transformer architectures.
Hands-on experience in Fine-tuning LLMs/S2S models for new language groups.
Expertise in CUDA 12.x and NVIDIA optimization libraries.
Fullstack Development:
Expert-level knowledge of WebRTC / WebSockets for real-time media streaming.
Demonstrated experience in developing Telegram Mini Apps (TMA).
Professional mastery of FastAPI and React / Next.js.
Strong understanding of the constraints and requirements of Low-latency systems.