Lip-Sync & Voice Quality Fix

Job ID: 40566637

Budget: ₹1,500 – ₹12,500 INR

I have an almost-finished video-localisation app: it already turns any user video into multiple languages by chaining speech-to-text, translation and text-to-speech. The last hurdle is the “believability” of the final clip. Right now both lip-movement timing and mouth shapes drift from the generated speech, and the audio itself feels flat because the recordings that feed the TTS are noisy. I want the final export to look and sound as if it were natively shot in the target language, with noticeably better modulation and intonation.

What I expect from you
• Diagnose and correct the lip-sync engine so phoneme-level mouth shapes and overall timing line up perfectly with the new audio.
• Design or fine-tune a voice pipeline (clean-up, enhancement or a different TTS model) that produces clearer speech with richer modulation and natural-sounding intonation.
• Hand back commented code or a plug-in that drops straight into my existing Python workflow (currently built around Wav2Lip, ffmpeg and a basic Tacotron-style TTS).

Acceptance will be a side-by-side comparison clip where the sync is visually convincing and the voice feels studio-quality without obvious artefacts.

Please only apply if you are an individual freelancer; I’m not considering agencies for this job.