Low Resource TTS Transformer Training

Job ID: 40065263

Budget: $3,000 – $5,000 USD

I have a very limited corpus of speech and transcripts—just a few hours—and I need it to sing in a brand-new language. The goal is to adapt a large ASR/TTS Transformer (Whisper-style architecture) to this low-data setting, squeezing the most out of the dataset through smart training techniques.

What matters most is the training strategy: curriculum scheduling, learning-rate tricks, mixed-precision, and any other techniques you know that keep a Transformer stable when the data are scarce. Data augmentation, synthetic data creation, few-shot and self-supervised learning are all on the table; noise injection, pitch or speed variation can be explored if they prove helpful.

I prefer to work within proven open-source stacks—Coqui TTS, Mozilla TTS, and Tacotron—so please build and document the pipeline in those environments. The final model should achieve intelligible, natural-sounding speech in the target language and be reproducible on a single high-end GPU.

Deliverables
• Cleaned, ready-to-train audio/text dataset (scripts included)
• Training code and configuration files for the chosen framework(s)
• Fine-tuned model checkpoints plus inference script
• Short report summarising hyper-parameters, augmentation methods, and evaluation results (WER/MOS)

Acceptance criteria
• Training scripts run end-to-end from raw data to synthesis without errors
• Objective metrics meet or exceed baseline Whisper fine-tune on the same data
• Synthesised samples judged 4.0 MOS or higher by at least three native speakers

If you enjoy pushing Transformers into low-resource territory, let’s make this language heard.