Multilingual Speech Emotion Recognition Specialist
Budget: ₹600 – ₹1,500 INR
I'm seeking a proficient Audio Machine Learning Engineer. The main objective of this project is to build, test, and implement a machine learning model that works effectively with multiple languages, interpreting user-uploaded audio files.
Work to carry out:
Make a unsupervised model of SER- Speech Emotion Recognition for 4 emotion i.e sad, happy,anger and fear.
- Dataset already created of 150 sample in .wav files at 16khz, 16bit mono audio with emotion
annotated of file name 4th letter. (Cannot share Dataset, as it is under process for research paper)
- Feature extraction: (make sperate folder after feature extracted)
1. Time-domain features - *short-term energy of signal, zero crossing rate, maximum amplitude, minimum energy, entropy of energy*
2. Frequency-domain features- * MFCC - [12 (form 26 mel freq), 13 delta,13 delta acceleration ], spectral centroid, spectral rolloff, spectral entropy and chroma coefficients*
- Do K-means clustering, make confusion matrix and do MLP perceptron. choose epochs as per
needed
- Make a user interface model where input audio of unknown sample (3- second) is given and it
can give o/p as emotion. (like use django or any other)
Key Responsibilities:
- Developing and implementing a speech recognition model
- Ensuring the model can accurately process multiple languages
Ideal Skills:
- Expertise in ML algorithms and audio datasets
- Experience with speech-to-text technologies and user-uploaded audio files
- Fluency in several languages would be beneficial
Applicants should have previous proven experience in similar projects and show a strong understanding of multilingual speech recognition. The goal is not just to recognize speech, but to convert it into useful, usable input for our system.
Payment - 100% (25 + 25 + 50)
Work to carry out:
Make a unsupervised model of SER- Speech Emotion Recognition for 4 emotion i.e sad, happy,anger and fear.
- Dataset already created of 150 sample in .wav files at 16khz, 16bit mono audio with emotion
annotated of file name 4th letter. (Cannot share Dataset, as it is under process for research paper)
- Feature extraction: (make sperate folder after feature extracted)
1. Time-domain features - *short-term energy of signal, zero crossing rate, maximum amplitude, minimum energy, entropy of energy*
2. Frequency-domain features- * MFCC - [12 (form 26 mel freq), 13 delta,13 delta acceleration ], spectral centroid, spectral rolloff, spectral entropy and chroma coefficients*
- Do K-means clustering, make confusion matrix and do MLP perceptron. choose epochs as per
needed
- Make a user interface model where input audio of unknown sample (3- second) is given and it
can give o/p as emotion. (like use django or any other)
Key Responsibilities:
- Developing and implementing a speech recognition model
- Ensuring the model can accurately process multiple languages
Ideal Skills:
- Expertise in ML algorithms and audio datasets
- Experience with speech-to-text technologies and user-uploaded audio files
- Fluency in several languages would be beneficial
Applicants should have previous proven experience in similar projects and show a strong understanding of multilingual speech recognition. The goal is not just to recognize speech, but to convert it into useful, usable input for our system.
Payment - 100% (25 + 25 + 50)
Related categories:
Python
Algorithm
Machine Learning (ML)
Digital Signal Processing
Audio Engineering