Multimodal Emotion Recognition AI Development
Budget: $250 – $750 USD
I'm looking for a skilled AI developer to create a multimodal model for accurately classifying emotions, leveraging the IEMOCAP dataset.
The model ultimately needs to be developed using Python, as that's the language I'm comfortable with. Experience in creating multimodal AI systems, specifically for emotion classification, is crucial for this job. Also, familiarity working with the IEMOCAP dataset will be highly advantageous.
To cut it short,
1) the dataset sizs is 8K records, data (audio, text and spectogram images) are all processed and parsed.
2) the goal is to build a deep learning model using Pytorch (tensorflow is an option too) where we compare the results of each modality separately, vs Multimodal using early, join or late fusion
3) I have made good progress, but my network layers /structure is not the best. I am opened for pretrained models
4) We can use GANs to generate more spectogram images (if needed)
Additional details
Here's what the project involves:
- Constructing a robust Multimodal AI model:
* Capable of concurrently handling audio, visual, and textual data.
* Effectively utilizing the diverse data types of the IEMOCAP dataset.
- Enhancing Existing Techniques:
* Improvements upon current emotion classification technologies.
I am flexible with both pay per hour or fixed project, we can work together to achieve the task
The model ultimately needs to be developed using Python, as that's the language I'm comfortable with. Experience in creating multimodal AI systems, specifically for emotion classification, is crucial for this job. Also, familiarity working with the IEMOCAP dataset will be highly advantageous.
To cut it short,
1) the dataset sizs is 8K records, data (audio, text and spectogram images) are all processed and parsed.
2) the goal is to build a deep learning model using Pytorch (tensorflow is an option too) where we compare the results of each modality separately, vs Multimodal using early, join or late fusion
3) I have made good progress, but my network layers /structure is not the best. I am opened for pretrained models
4) We can use GANs to generate more spectogram images (if needed)
Additional details
Here's what the project involves:
- Constructing a robust Multimodal AI model:
* Capable of concurrently handling audio, visual, and textual data.
* Effectively utilizing the diverse data types of the IEMOCAP dataset.
- Enhancing Existing Techniques:
* Improvements upon current emotion classification technologies.
I am flexible with both pay per hour or fixed project, we can work together to achieve the task