Natural Voice TTS for Applications
Budget: $30 – $250 USD
We are seeking an experienced developer or team to create an open-source AI Text-To-Speech model that emulates the capabilities of Eleven Labs.
I will suggest to use Coqui XTTS model and integrate it with HifiGANs and other models to train and fine tune it.
Or
Use VITS + GST + HiFi-GAN + Neural Codec
The latency should be less than 0.7 seconds.
The sound should be exactly human like with fluid flow of expression, emotion and tone.
The ideal candidate will have a strong background in natural language processing/NLP and machine learning, with a focus on speech synthesis. Your task will involve designing, training, and optimizing the model to produce high-quality, human-like speech. Collaboration and open-source principles are essential, as we aim to provide this tool to the community. If you have a passion for AI and voice technology, we want to hear from you!
I will suggest to use Coqui XTTS model and integrate it with HifiGANs and other models to train and fine tune it.
Or
Use VITS + GST + HiFi-GAN + Neural Codec
The latency should be less than 0.7 seconds.
The sound should be exactly human like with fluid flow of expression, emotion and tone.
The ideal candidate will have a strong background in natural language processing/NLP and machine learning, with a focus on speech synthesis. Your task will involve designing, training, and optimizing the model to produce high-quality, human-like speech. Collaboration and open-source principles are essential, as we aim to provide this tool to the community. If you have a passion for AI and voice technology, we want to hear from you!