Building a Text-to-Speech system using Deep Learning speech synthesis model

Job ID: 32830672

Budget: $250 – $750 USD

I want to build a speech synthesis system (TTS) for the Kurdish language. I have collected <audio,text> pair. I need to configure the environment and build the deep learning model to train on the data I have collected. There are two projects that can be used, please find the link to both projects below. And I'm open to new suggestions as well.

https://github.com/NVIDIA/tacotron2
https://espnet.github.io/espnet/