DL ML training on voice commands -- 2

Job ID: 36424820

Budget: $250 – $750 USD

I am looking for an experienced developer to build server application to do real-time fine tuning for Whisper ASR model(tiny), which will provide REST API to the user’s device to train a small set of voice command data for farm automation control. This project will include
1. optimizing whisper’s pretrained data to Korean and Farm automation command set.
2. real time(less than 30 seconds) fine tuning of voice common(less than 5 seconds voice) under AMD Ryzen 9 7950X and GTX4090 system.
3. REST API to user’s device ( upload training data and downloading trained model)

if you don't have necessary hardware setup, you can access my on-premise server or we will set up a cloud server for this. The multi-language dataset for Korean is not provided in huggingface but you can download one from this link which need data validation;
https://aihub.or.kr/aihubdata/data/view.do?currMenu=115&topMenu=100&aihubDataSe=realm&dataSetSn=568
The hugginface fine-tunning blog:
https://github.com/huggingface/blog/blob/main/fine-tune-whisper.md
The whisper blog;
https://openai.com/research/whisper
The deliverables should be available via github and implemented on the afore-mentioned server, with a documentation.
Related categories: Python VoiceXML Tensorflow Deep Learning GitHub