Machine Learning Project - Censorship, Data Quality, Data Extraction, Data Analyzation. -- 2

Job ID: 31080684

Budget: $250 – $750 USD

*IMPORTANT* - You will be working with another app developer. The app developer will set up the backend if needed. Models do not need to be built from scratch, you can use any services that that do the requested work if they exist.

Overview:
I am building a talent scouting app that needs some Machine Learning built. I am looking for facial recognition to be built for photo verification. I am looking for a video quality analyzer to be built, in order to ensure nothing inappropriate is said or done in videos/audios and to ensure video is in good quality (not fuzzy if possible). Analyzer for a song sung by the original creator and compare against the submitted video of someone singing. If they are hitting the notes, the video submission will be a top performer. Analyzer for an acting video submission for an emotion through ML. Eye contact, voice projection, etc will be analyzed and a score will be given to the video.

Specifics:
For Video Analyzer Quality: No curse words should be allowed, so the analyzer should screen for any curse words being said and bleep them. No fuzzy video should be allowed. The video should at least be decent quality (480p). If quality is not good, then video should be kicked back. Also, person should be fully clothed in video, so ML should recognize shirts, pants, dresses, shorts, etc.

For song recognition: The ML should recognize if a song from a submission is not an original song. If it is not, then the ML should pick up which notes are in the original song and compare those song notes to the song notes in the submission. If the submission is hitting the originals notes consistently, they will be given a score. Highest score will be on recommended page.

Analyze acting submission: ML should recognize eye contact, analyze voice projection and speaking clearness, and then put together a score. The person with the highest score will be on the recommended page. Voice projection means analyze the decibels. If the decibals are under a certain threshold (quieter voice) then they will get a lower score.

Summary:
1. censorship - bleep curse word (data extraction first, from image or from audio) 2. censorship - identify porno 3. data quality - identify bad quality video 4. data extraction - identify eye contact and do calculation 5. data extraction - volumn(or other audio index) calculation Audio 1. data extraction - extract lyrics from submitted song 2. data analyze - do matching and give a score

Extra:
The censorship related content will be hard to build from scratch, so I would suggest to use Ali Cloud. CN is good at censorship: https://www.alibabacloud.com/help/doc-detail/53425.html?spm=a2c5t.10695662.1996646101.searchclickresult.552830ddRdFFBg