Optimization of Multi-Voice Text-to-Speech (TTS) Models (South Park Like cartman, KYLE, STAN etc) READ FIRST!

Job ID: 37583029

Budget: $30 – $250 USD

You will be working together with us and you will need to explain and show us how to do these things as well

We are seeking an experienced Audio Engineer or Machine Learning Specialist to assist in the fine-tuning and optimization of our pre-trained Text-to-Speech (TTS) models. Currently, we have five distinct voices and plan to expand to more. Our aim is to refine these models to achieve a high level of consistency and naturalness across all voices.

Objectives:

To perform detailed analysis and optimization of existing TTS models.
To ensure that all voices maintain consistent quality and sound natural.
To adjust and fine-tune the voice models to match a specified reference tone and style.
Scope of Work:

Analyze the current performance of our TTS models.
Identify areas for improvement in prosody, intonation, and clarity.
Implement adjustments to voice models to enhance quality and consistency.
Conduct A/B testing with sample voice outputs to determine successful optimizations.
Project Deliverables:

A report detailing the current state of the TTS models and identified areas for improvement.
Optimized TTS models with clear documentation of the changes made.
Audio samples demonstrating the improvements in voice quality and consistency.
Required Skills:

Strong background in audio engineering or machine learning, specifically in the realm of voice or speech.
Experience with TTS systems and understanding of DSP (Digital Signal Processing).
Proficiency with machine learning tools and libraries relevant to TTS (e.g., TensorFlow, PyTorch).
Ability to work with audio analysis and editing software.