Building concatenative TTS

Job ID: 33741915

Budget: $250 – $750 USD

I need to build a Text-to-speech (speech synthesis) system using Festival framework through unit selection TTS approach or diaphine concatenation approach. I have a dataset of 6 hours of recordings that’s segmented to 0-10 seconds length, and a separate library for diaphone units. Attached is a document that can be used as a reference to implement this system.

In addition to the system a detailed report should be provided that includes (introduction, methodology, how it was implemented, results and testing the results, conclusion)