Finetuning Whisper AI Speech to Text
Budget: €30 – €250 EUR
Hi Guys,
we are using whisper to transcribe Audofiles from Diskussions with multiple Persons to Text. Whisper works very well so far. The problem is that the discussion is mostly in healthcare and there are some special words that the german model is not trained for, so whisper doesnt recognize them right. Im using the general large modelv2. On huggingface there are some german trained models to use for finetuning may this can gelp a little, but at end i think i need traing an own model. We have enough audio files with 100% correct transcriptions to make training with them wich contains these words that whisper should learn to recognize in future right. I need support to make this training and create own model and use them for our audio transcriptions.
Introducing whisper:
https://openai.com/research/whisper
Open-Source:
https://github.com/openai/whisper
Whisper finetuning tutorial:
https://huggingface.co/blog/fine-tune-whisper
HuggingFace German model to try:
https://huggingface.co/bofenghuang/whisper-large-v2-cv11-german
So the final goal of this project is that whisper recognize the german speech better then before because the more specific trained data, and also recognize our special healthcare words right. If you arent familar with this taks please save our both time. Because you will just get the 100% payout when thats accomplished. And if you cant do in time, or cant provide the quality i will dispute and hire someone else.
we are using whisper to transcribe Audofiles from Diskussions with multiple Persons to Text. Whisper works very well so far. The problem is that the discussion is mostly in healthcare and there are some special words that the german model is not trained for, so whisper doesnt recognize them right. Im using the general large modelv2. On huggingface there are some german trained models to use for finetuning may this can gelp a little, but at end i think i need traing an own model. We have enough audio files with 100% correct transcriptions to make training with them wich contains these words that whisper should learn to recognize in future right. I need support to make this training and create own model and use them for our audio transcriptions.
Introducing whisper:
https://openai.com/research/whisper
Open-Source:
https://github.com/openai/whisper
Whisper finetuning tutorial:
https://huggingface.co/blog/fine-tune-whisper
HuggingFace German model to try:
https://huggingface.co/bofenghuang/whisper-large-v2-cv11-german
So the final goal of this project is that whisper recognize the german speech better then before because the more specific trained data, and also recognize our special healthcare words right. If you arent familar with this taks please save our both time. Because you will just get the 100% payout when thats accomplished. And if you cant do in time, or cant provide the quality i will dispute and hire someone else.