Realtime prediction and voice recognition

Job ID: 36434754

Budget: $250 – $750 USD

This project will include
1. optimizing whisper’s pretrained data to Korean and Farm automation command set.
2. real time(less than 30 seconds) fine tuning of voice common(less than 5 seconds voice) under AMD Ryzen 9 7950X and GTX4090 system.
3. REST API to user’s device ( upload training data and downloading trained model)

if you don't have necessary hardware setup, you can access my on-premise server or we will set up a cloud server for this. The multi-language dataset for Korean is not provided in huggingface but you can download one from this link which need data validation;
https://aihub.or.kr/aihubdata/data/view.do?currMenu=115&topMenu=100&aihubDataSe=realm&dataSetSn=568
or download from here (tyeng0418); jupyter1.d.tyeng.com
The hugginface fine-tunning blog:
https://github.com/huggingface/blog/blob/main/fine-tune-whisper.md
The whisper blog;
https://openai.com/research/whisper
The deliverables should be available via github and implemented on the afore-mentioned server, with a documentation.

4 milestones;
1- developing the flask app
2- training data of korean whispers data and optimize it.
3- make real time prediction
4- REST API for any device
Related categories: Python Big Data Sales Tensorflow Deep Learning GitHub