AI/ML Expert Needed for TTS Model Enhancement
Budget: €750 – €1,500 EUR
We are seeking a highly skilled AI/ML freelancer to fine-tune a multilingual Text-to-Speech (TTS) model capable of generating ultra-realistic human voices. The goal is to achieve a voice quality that is 85% or more of the quality provided by ElevenLabs, incorporating natural human-like expressions such as breathing, whispering, laughing, emotional tone, and natural prosody.
Project Requirements:
Fine-tune an existing TTS model (or propose a custom approach) to enhance realism.
Implement breath sounds, whispering, laughter, emotional variations, and natural pauses to mimic real human speech.
Optimize for natural prosody, cadence, and intonation across different emotions and speaking styles.
Ensure smooth phoneme transitions to eliminate robotic or unnatural artifacts.
Support multiple languages with native-level fluency and accuracy.
Deliver a model that can generate high-quality speech in real-time or near-real-time.
Ideal Candidate:
Proven experience in TTS fine-tuning, deep learning, and neural speech synthesis.
Familiarity with state-of-the-art models like ElevenLabs, VALL-E, Tacotron, FastSpeech, or similar.
Expertise in audio processing, phonetics, multilingual speech synthesis, and language modeling.
Strong knowledge of PyTorch/TensorFlow, speech datasets, and audio augmentation techniques.
Ability to demonstrate past projects or a portfolio related to TTS.
Deliverables:
A trained and fine-tuned multilingual TTS model meeting the quality requirements.
A demo showcasing the model’s capabilities (preferably with samples covering multiple languages, emotions, and speaking styles).
Documentation on model architecture, training process, language integration, and usage instructions.
If you have the expertise to push multilingual TTS technology to the next level, we’d love to hear from you! Please provide examples of past work and your proposed approach for this project.
Budget: Open for discussion based on expertise and project scope.
Timeline: To be discussed.
Looking forward to your proposals!
Project Requirements:
Fine-tune an existing TTS model (or propose a custom approach) to enhance realism.
Implement breath sounds, whispering, laughter, emotional variations, and natural pauses to mimic real human speech.
Optimize for natural prosody, cadence, and intonation across different emotions and speaking styles.
Ensure smooth phoneme transitions to eliminate robotic or unnatural artifacts.
Support multiple languages with native-level fluency and accuracy.
Deliver a model that can generate high-quality speech in real-time or near-real-time.
Ideal Candidate:
Proven experience in TTS fine-tuning, deep learning, and neural speech synthesis.
Familiarity with state-of-the-art models like ElevenLabs, VALL-E, Tacotron, FastSpeech, or similar.
Expertise in audio processing, phonetics, multilingual speech synthesis, and language modeling.
Strong knowledge of PyTorch/TensorFlow, speech datasets, and audio augmentation techniques.
Ability to demonstrate past projects or a portfolio related to TTS.
Deliverables:
A trained and fine-tuned multilingual TTS model meeting the quality requirements.
A demo showcasing the model’s capabilities (preferably with samples covering multiple languages, emotions, and speaking styles).
Documentation on model architecture, training process, language integration, and usage instructions.
If you have the expertise to push multilingual TTS technology to the next level, we’d love to hear from you! Please provide examples of past work and your proposed approach for this project.
Budget: Open for discussion based on expertise and project scope.
Timeline: To be discussed.
Looking forward to your proposals!