Real-Time AI Avatar Lip-Sync Specialist

Job ID: 40330568

Budget: €250 – €750 EUR

We’re building a real-time conversational AI avatar platform and need a specialist to train and deploy a neural lip-sync model on a single avatar persona. We already have reference video + audio — this is not an R&D project, we need someone who’s done this before and can execute fast.
The goal: audio-in → lip-synced video frames out, real-time (-100ms/frame), ≤4 GB VRAM per session so we can scale to 15–20 concurrent sessions on a single H100 80GB.
*This is Phase 1 (proof of concept on 1 persona). If results are good, Phase 2 is a paid follow-up to build the full automation pipeline (multi-persona training, audio generation, WebRTC streaming integration).*


Phase 1
Model Selection & Benchmark
Pick the best base model for our constraints ( GeneFace++ — or propose something better)
∙ Quick benchmark on our H100: quality vs. latency vs. VRAM
∙ Validate French phoneme handling
Milestone 2: Fine-Tune + Optimize (Days 4–10)
∙ Fine-tune on our avatar footage — identity-locked, artifact-free output
∙ Optimize inference: TensorRT/ONNX, FP16/INT8 quantization
∙ Target: ≤l-4 GB VRAM, -100ms latency per frame
∙ Deliver a working FastAPI endpoint: audio stream in → video frames out
∙ Docker container, reproducible
Deliverables
1. Trained model checkpoint for our persona
2. Inference server (FastAPI) in Docker
3. Benchmark numbers (latency, VRAM, visual quality)
4. Brief documentation


Phase 2
Phase 2 — Follow-Up Contract (If Phase 1 Succeeds)
∙ Automated pipeline to train new personas from raw footage
∙ Full integration with Pipecat + WebRTC streaming
∙ Multi-session scaling & load management
∙ Audio generation pipeline integration
∙ Budget and scope discussed separately