Speech to Text and Voice Biometrics

Job ID: 34576950

Budget: $50 – $0 CAD

My team is working on a project that uses Nvidia Nemo to convert speech to text (Russian
and English)and client voice identification. We need to create a Russian model. Nvidia's
Russian model is not accurate enough. We have following initial questions.

1) Is it possible to use one model to dynamically diarize multiple audio files
2) How to augment a model using last check point
3) How many epochs needed to create a model
4) We need to identify speaker language (Russian or English) and use correct model
5) Is it possible to create one model for both Russian and English
6) Hardware spec requirements for creating a ASR model
Related categories: Python Data Science Pytorch