NLP & Data Modeling
Budget: ₹600 – ₹3,000 INR
You are provided with a large number of Wikipedia comments which have been labeled by human raters for toxic behavior. The types of toxicity are:
toxic
severe_toxic
obscene
threat
insult
identity_hate
Task:
Create models that predict the probability of each type of toxicity for each comment.
Perform a descriptive statistical analysis and make interesting inferences.
Design and explain the data science pipeline followed in this task involving Data preprocessing, Data Analysis, Feature Engineering, Modeling, and Evaluation.
Data can be found here:
https://www.kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge/data
toxic
severe_toxic
obscene
threat
insult
identity_hate
Task:
Create models that predict the probability of each type of toxicity for each comment.
Perform a descriptive statistical analysis and make interesting inferences.
Design and explain the data science pipeline followed in this task involving Data preprocessing, Data Analysis, Feature Engineering, Modeling, and Evaluation.
Data can be found here:
https://www.kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge/data