Natural Voice TTS for Applications

Job ID: 39141383

Budget: $30 – $250 USD

We are seeking an experienced developer or team to create an open-source AI Text-To-Speech model that emulates the capabilities of Eleven Labs.

I will suggest to use Coqui XTTS model and integrate it with HifiGANs and other models to train and fine tune it.

Or

Use VITS + GST + HiFi-GAN + Neural Codec

The latency should be less than 0.7 seconds.
The sound should be exactly human like with fluid flow of expression, emotion and tone.

The ideal candidate will have a strong background in natural language processing/NLP and machine learning, with a focus on speech synthesis. Your task will involve designing, training, and optimizing the model to produce high-quality, human-like speech. Collaboration and open-source principles are essential, as we aim to provide this tool to the community. If you have a passion for AI and voice technology, we want to hear from you!