Unreal Engine Real-Time MetaHuman Voice Talkbot

Job ID: 40258956

Budget: $250 – $750 USD

I need a developer who can wire together Unreal Engine’s MetaHuman framework with several AI services to create a real-time, push-to-talk assistant that speaks back instantly in perfect lip-sync.

Two separate screens drive the experience.
• On the landscape display (1920×1080) I’ll host English and Arabic buttons, a press-and-hold mic icon, plus live text showing what the user just said and what the AI replies.
• On the portrait display (1080×1920) a MetaHuman avatar delivers the reply, streaming audio while its face tracks the phonemes.

My chosen pipeline is:
1. Button held → audio captured.
2. Speech sent to the OpenAI Whisper API for transcription (I prefer the cloud API or local model).
3. Plain text routed through n8n, which handles prompt logic and returns the response.
4. That response feeds straight into ElevenLabs for low-latency audio streaming.
5. MetaHuman receives the stream and plays it with accurate lip-sync.

Top priority is minimal latency; every millisecond matters along with lip-sync. The language switch happens only through the on-screen toggle—no voice commands or auto-detection.

Deliverables I expect:
• Full Unreal Engine project with clean, well-commented Blueprint code and neatly organised folders
• A packaged build that runs out of the box on Windows
• Setup notes explaining API keys, n8n endpoint configuration, and any special Unreal plugins or MetaHuman Live Link steps

When you reply, highlight past work that combines MetaHuman (or similar real-time avatars) with external AI services, especially anything that shows you’ve already tuned pipelines for sub-second round-trips. Please include a realistic timeline from project kick-off to first test build and to final delivery.