Deep Fake Face Animation, Simulation and Lip Sync

Job ID: 31353867

Budget: $750 – $1,500 USD

Alef company has already set up AWS ec2 servers in Singapore location and China Beijing location for its multilingual speech recognition, text to speech, and translation http web services. The Alef APIs offered to our clients are json files through the endpoints located in AWS Singapore and China.

This project is about Face animation and Talking head lip sync and 3D simulation.

Deepfake uses motion tracking, key points, images, facial recognition, neural networks.

https://thispersondoesnotexist.com/

We can have a meeting in order to clarify the project and sign NDA etc. We need an emotion predictor in this project in addition to lip synch when talking head speaking.

please take a look at these links for lip sync, they used wav2lip + GAN
it is open source code at Github

https://www.theverge.com/21428653/lip-sync-ai-deepfake-wav2lip-code-how-to

https://medium.com/deepgamingai/deepfakes-ai-improved-lip-sync-animations-with-wav2lip-b5d4f590dcf

https://www.theverge.com/2019/5/23/18637373/deepfakes-samsung-ai-research-results-single-photo-algorithm

few shot learning is also possible

https://www.unite.ai/googles-lipsync3d-offers-improved-deepfaked-mouth-movement-synchronization/

if your proposal can include the latest method and new technology will be great!

here is another example Flawless website

https://gizmodo.com/deepfake-lips-are-coming-to-dubbed-films-1846840191



1. phoneme_timestamps object is useful for animation purposes. An animated model can easily synchronize mouth movements with the synthesized audio.

2. Our video recording of singer has a small angle due to recording camera.

in the sample file, you can find the text grid with Chinese pinyin like xiang3, and its timestamp and corresponding wav.

More resources are following:

https://nv-adlr.github.io/view-generalization

https://github.com/yunjey/stargan

GAN for style-transfer, fake face, face animation, starGAN

- Diverse Image Synthesis for Multiple Domains- StarGAN2
A good image-to-image translation model should learn a mapping between different visual domains while satisfying the following properties
- AnimeGAN: A Novel Lightweight GAN for Photo Animation
- Unsupervised Generative Attentional Networks with Adaptive Layer-Instance Normalization for Image-to-Image Translation(minivision)
- Super resolution using Sharpen AI