High-Accuracy Face Lipsync Swap
Budget: $10 – $30 USD
I need a reliable pipeline that takes any pre-recorded video, swaps the on-screen face with a chosen target face, and keeps the lips perfectly synced to the original soundtrack. High accuracy is non-negotiable—phoneme-level alignment, natural mouth shapes, and seamless facial blending should hold up under close inspection and slow-motion playback.
Preferred stack
You’re free to combine proven solutions such as Wav2Lip, DeepFaceLab, FaceSwap, or custom GAN models, as long as the end result meets the visual standard. Python with PyTorch/TensorFlow is ideal because I want the option to retrain or fine-tune later, but an all-in-one executable is acceptable if thoroughly documented.
Required deliverables
• Use e.g. Kling AI, Wan 2.3 Runway ML
Acceptance criteria
• Mouth movements match the audio with no visible lag (≤1 frame tolerance)
• Skin tone and lighting blend convincingly; no flicker or jitter across frames
• Output video keeps the original resolution and frame rate
• Solution runs on a single high-end GPU (e.g., RTX 3090) without crashing
Once everything works as specified, I’ll test on a separate video set to confirm generalisation before sign-off.
Preferred stack
You’re free to combine proven solutions such as Wav2Lip, DeepFaceLab, FaceSwap, or custom GAN models, as long as the end result meets the visual standard. Python with PyTorch/TensorFlow is ideal because I want the option to retrain or fine-tune later, but an all-in-one executable is acceptable if thoroughly documented.
Required deliverables
• Use e.g. Kling AI, Wan 2.3 Runway ML
Acceptance criteria
• Mouth movements match the audio with no visible lag (≤1 frame tolerance)
• Skin tone and lighting blend convincingly; no flicker or jitter across frames
• Output video keeps the original resolution and frame rate
• Solution runs on a single high-end GPU (e.g., RTX 3090) without crashing
Once everything works as specified, I’ll test on a separate video set to confirm generalisation before sign-off.