Speech to Text and Voice Biometrics
Budget: $50 – $0 CAD
My team is working on a project that uses Nvidia Nemo to convert speech to text (Russian
and English)and client voice identification. We need to create a Russian model. Nvidia's
Russian model is not accurate enough. We have following initial questions.
1) Is it possible to use one model to dynamically diarize multiple audio files
2) How to augment a model using last check point
3) How many epochs needed to create a model
4) We need to identify speaker language (Russian or English) and use correct model
5) Is it possible to create one model for both Russian and English
6) Hardware spec requirements for creating a ASR model
and English)and client voice identification. We need to create a Russian model. Nvidia's
Russian model is not accurate enough. We have following initial questions.
1) Is it possible to use one model to dynamically diarize multiple audio files
2) How to augment a model using last check point
3) How many epochs needed to create a model
4) We need to identify speaker language (Russian or English) and use correct model
5) Is it possible to create one model for both Russian and English
6) Hardware spec requirements for creating a ASR model