Real-Time AI Avatar Streaming
Budget: €30 – €250 EUR
I want to take my livestreams on macOS to the next level with a real-time AI avatar that feels alive on camera and is ready to drop straight into OBS as a virtual source. Here is the core of what I need built:
• The avatar must track my face through an ordinary webcam, translate subtle expressions in real time, and lip-sync perfectly to whatever I say on-mic.
• I need the voice layer to offer several high-quality, fully licensed synthetic voices that I can switch on the fly. All voice cloning or TTS work has to be ethical and consent-based; no grey-area datasets.
• Latency has to stay low enough for professional broadcasting—ideally under 50 ms from camera to virtual camera output.
• Everything has to run smoothly on Apple-silicon Macs and slot into OBS without extra hoops.
I have outlined additional context, performance targets, and reference links here:
https://docs.google.com/document/d/1Tub_ZZnYJBMihfL4JJqYXwS8gtWk8UM953i1UBejALM/edit?tab=t.0
Deliverables
1. A macOS application or hardened OBS plug-in that:
– Reads a live webcam feed
– Performs facial recognition & expression mapping
– Generates a synced avatar output as a virtual camera
– Integrates multiple selectable synthetic voices
2. Source code and build instructions (M1/M2 compatible).
3. A quick-start guide so I can test, swap voices, and push the signal live within OBS.
Acceptance criteria
• Avatar expression and lip-sync accuracy visually match a side-by-side real feed.
• Switching voices mid-stream introduces no audible artifacts or drift.
• End-to-end latency ≤ 50 ms on an M2 Pro.
• No personal biometric or voice data leaves the local machine.
If you have shipped similar real-time ML or avatar projects—especially using frameworks like MediaPipe, TensorFlow, PyTorch, Core ML, or voice APIs such as ElevenLabs—let’s talk. I’m ready to move quickly once I see a clear technical plan and architecture outline.
• The avatar must track my face through an ordinary webcam, translate subtle expressions in real time, and lip-sync perfectly to whatever I say on-mic.
• I need the voice layer to offer several high-quality, fully licensed synthetic voices that I can switch on the fly. All voice cloning or TTS work has to be ethical and consent-based; no grey-area datasets.
• Latency has to stay low enough for professional broadcasting—ideally under 50 ms from camera to virtual camera output.
• Everything has to run smoothly on Apple-silicon Macs and slot into OBS without extra hoops.
I have outlined additional context, performance targets, and reference links here:
https://docs.google.com/document/d/1Tub_ZZnYJBMihfL4JJqYXwS8gtWk8UM953i1UBejALM/edit?tab=t.0
Deliverables
1. A macOS application or hardened OBS plug-in that:
– Reads a live webcam feed
– Performs facial recognition & expression mapping
– Generates a synced avatar output as a virtual camera
– Integrates multiple selectable synthetic voices
2. Source code and build instructions (M1/M2 compatible).
3. A quick-start guide so I can test, swap voices, and push the signal live within OBS.
Acceptance criteria
• Avatar expression and lip-sync accuracy visually match a side-by-side real feed.
• Switching voices mid-stream introduces no audible artifacts or drift.
• End-to-end latency ≤ 50 ms on an M2 Pro.
• No personal biometric or voice data leaves the local machine.
If you have shipped similar real-time ML or avatar projects—especially using frameworks like MediaPipe, TensorFlow, PyTorch, Core ML, or voice APIs such as ElevenLabs—let’s talk. I’m ready to move quickly once I see a clear technical plan and architecture outline.