Arif project

Job ID: 34336549

Budget: $10 – $30 USD

Arif will finish two tasks

1) We want to run Blenderbot 3.0 on a cloud server as an api to communicate with. You need to add the proper model like Qzoo:blenderbot2/blenderbot2_400M/model. You can send it to us as a docker image, we build it on our cloud, communicate with it through the api you provide and once all is working properly, this task is tagged as done.
- We want you to video record the process of doing so as we are interested mainly to learn how to do it so we apply it on our servers and tag this task as complete.
-----------------------------------------------------------------------------------------------------
2) build a ready to use docker for SpeechBrain platform, including the models required to make its functions work best.

Here are speechbrain key links (it's for you to check if there is anything else needed):

https://speechbrain.github.io/
https://github.com/speechbrain/speechbrain
https://huggingface.co/speechbrain

What we expect:

- Through the docker, we expect all the functions included in the platform to be working properly, including the Key features listed in the github link, based on the proper models available (check the links).

- We expect to get a clear path to communicate with each function (Curl, api, etc...). For example, for the STT, we will test it by posting an audio and checking the text we get. For the TTS, we will post a text and get audios with different voice tones (as the platform includes the possibility to choose the voice). etc...

After you are done, we will try it all locally on our device and on the cloud, and if all is good, the project will be tagged as done, then we move with you to another project.

The main functions required (check their details in the above GitHub link):

We start with the main ones:


– Speech recognition, Check this video from where I am sharing it:https://www.youtube.com/watch?v=TfgnmfxsPXY&t=2352s

So, basically I will create a desktop exe with a button in it, click it to record, once done speaking, I get back the audio file, the text of it through the STT, with the spoken language info.
– Text to speech:
– Grapheme-to-Phoneme (G2P)
– Translation. I set the language to translate to and get a translation to my input.

Speech recognition captures speech and post it back as an audio file and as a text.

TTS I send text to and it responds with an audio based on the selected voice tone.

G2p

Translation

We want you to video record the process of doing so as we are interested mainly to learn how to do it so we apply it on our servers and tag this task as complete.
-----------------------------------------------------------------------------

Arif will deliver the full projects for testing within 2 days from when it is rewarded or else we are free to cancel the milestones with immediate effect. After Arif deliver the two tasks we will test them within 48 hours and if all is good we release the milestones.

Milestones will be released only if ALL the requirements are fulfilled and on time. Partly-done equals not done.

Best luck!
Related categories: Python Linux Docker Pytorch Server