Windows TTS Voice Fine-Tuning

Job ID: 40040212

Budget: €30 – €250 EUR

I have a Windows-based desktop application that already speaks, but I want it to sound unmistakably natural in Español Castellano and allow users to switch to this new custom voice on demand. Your task is to fine-tune an existing neural text-to-speech model so it delivers that polished Castilian tone, then package it so my app can call it locally without relying on any cloud service.

You’re free to work with the toolkit you know best—Tacotron-2, FastSpeech-2, Glow-TTS, VITS, or a comparable neural approach—as long as the final runtime is lightweight, runs offline on Windows 10/11, and exposes a simple DLL or command-line interface my app can invoke.

Deliverables
• Fine-tuned Castilian Spanish voice model files
• Windows-ready inference runtime (ONNX, Piper, or similar)
• Demo CLI that converts arbitrary text to WAV
• Integration guide covering API calls and voice-switch steps

Acceptance criteria
• Naturalness: MOS ≥ 4.0 on five sample sentences
• Latency: < 300 ms for a 100-character input on mid-range CPU
• Fully offline execution once installed

If you have previous work that showcases custom voices on Windows, feel free to include a short demo clip or repo link along with your proposed approach.