Lifelike AI-Generated Podcast Video Co-Hosts
Budget: €8 – €30 EUR
? AI Podcast Video Generation Brief
? Project Title:
Creating AI Podcasts with Custom Video Avatars – Male/Female Co-Host Version
? Purpose:
Create a high-quality, lifelike AI podcast video with two speakers (male + female), using real photos of each person as the visual base. The goal is to synchronize their speech with the provided joint audio track, and generate realistic mouth movements, expressions, gestures, and natural mannerisms that match the energy, pauses, and tone of their voices.
?? Speaker 1 (Male)
Role: Host / Co-narrator
Audio: Male parts from uploaded .wav file
Image File: a-medium-shot-of-an-attractive-man...jpeg
Expression Guidelines: Calm, curious, thoughtful — occasionally raising eyebrows or nodding when reacting to surprising insights
Mouth Sync: Align precisely with his speaking parts (Speaker 1 from transcript)
?? Speaker 2 (Female)
Role: Co-host / Explainer
Audio: Female parts from uploaded .wav file
Image File: a-professional-portrait-photograph-of-an...jpeg
Expression Guidelines: Confident, engaging, friendly — smiling during positive insights, expressive hand gestures if supported
Mouth Sync: Align with Speaker 2 sections from transcript
?️ Files Provided:
?️ Audio File: Creating AI Podcasts with Custom Video Avatars.wav (joint audio of both speakers)
?♂️ Male Image: a-medium-shot-of-an-attractive-man...jpeg
?? Female Image: a-professional-portrait-photograph-of-an...jpeg
?️ Video Requirements:
?️ Format: Side-by-side or alternating camera angles for both speakers
? Emotionally aware facial movements, natural blinking, lip-sync precision
? Use Runway-style smoothing if available to reduce uncanny valley effect
? Optional: Add subtitles for accessibility (optional but preferred)
? Duration: Full length of audio (~[input duration here])
? Style: Clean background, minimal distraction — simulate a podcast studio vibe
? Output Format: MP4 or MOV, high-resolution (1080p or higher)
? Important Notes:
Use Speaker Diarization or timestamp mapping to split who says what
Do not auto-generate avatars — use the uploaded photos only
Emphasize realism over flashiness — this should look like a real podcast video, not a cartoon or overly animated avatar
I will send the two audios to you once the understanding of the assignment is established
? Project Title:
Creating AI Podcasts with Custom Video Avatars – Male/Female Co-Host Version
? Purpose:
Create a high-quality, lifelike AI podcast video with two speakers (male + female), using real photos of each person as the visual base. The goal is to synchronize their speech with the provided joint audio track, and generate realistic mouth movements, expressions, gestures, and natural mannerisms that match the energy, pauses, and tone of their voices.
?? Speaker 1 (Male)
Role: Host / Co-narrator
Audio: Male parts from uploaded .wav file
Image File: a-medium-shot-of-an-attractive-man...jpeg
Expression Guidelines: Calm, curious, thoughtful — occasionally raising eyebrows or nodding when reacting to surprising insights
Mouth Sync: Align precisely with his speaking parts (Speaker 1 from transcript)
?? Speaker 2 (Female)
Role: Co-host / Explainer
Audio: Female parts from uploaded .wav file
Image File: a-professional-portrait-photograph-of-an...jpeg
Expression Guidelines: Confident, engaging, friendly — smiling during positive insights, expressive hand gestures if supported
Mouth Sync: Align with Speaker 2 sections from transcript
?️ Files Provided:
?️ Audio File: Creating AI Podcasts with Custom Video Avatars.wav (joint audio of both speakers)
?♂️ Male Image: a-medium-shot-of-an-attractive-man...jpeg
?? Female Image: a-professional-portrait-photograph-of-an...jpeg
?️ Video Requirements:
?️ Format: Side-by-side or alternating camera angles for both speakers
? Emotionally aware facial movements, natural blinking, lip-sync precision
? Use Runway-style smoothing if available to reduce uncanny valley effect
? Optional: Add subtitles for accessibility (optional but preferred)
? Duration: Full length of audio (~[input duration here])
? Style: Clean background, minimal distraction — simulate a podcast studio vibe
? Output Format: MP4 or MOV, high-resolution (1080p or higher)
? Important Notes:
Use Speaker Diarization or timestamp mapping to split who says what
Do not auto-generate avatars — use the uploaded photos only
Emphasize realism over flashiness — this should look like a real podcast video, not a cartoon or overly animated avatar
I will send the two audios to you once the understanding of the assignment is established