Whisper speech to text Model Fine-Tuning for Bengali language

Job ID: 40020336

Budget: $30 – $250 USD

I like to work on the Whisper Small speech-to-text model and make it truly understand Bengali.we require someone who can handle the end-to-end technical work: cleaning and aligning our audio/text pairs, setting up the training pipeline in PyTorch, running the fine-tuning, and proving the quality gains with solid WER/CER evaluations.

Natural language generation techniques for data augmentation are welcome if they help squeeze more accuracy out of limited resources.

Please bring practical Whisper or Hugging Face training experience, a GPU-ready workflow, and a clear plan for reproducible experiments.

Deliverables
• Pre-processed and documented Bengali dataset ready for Whisper
• Training scripts/notebooks and configuration files
• Fine-tuned Whisper Small checkpoint(s)
• Evaluation report comparing base vs. tuned model on held-out test audio
• Simple inference script to demo the new model locally or via API

I’m ready to start as soon as you can outline your approach and estimated timeline. Our air to build a Bangeli Human voice agent. and its its a first step to our goal.