NLP and Data Modeling

Job ID: 34325908

Budget: ₹600 – ₹3,000 INR

You are provided with a large number of Wikipedia comments which have been labeled by human raters for toxic behavior. The types of toxicity are:

toxic
severe_toxic
obscene
threat
insult
identity_hate

Task:
Create models that predict the probability of each type of toxicity for each comment.
Perform a descriptive statistical analysis and make interesting inferences.
Design and explain the data science pipeline followed in this task involving Data preprocessing, Data Analysis, Feature Engineering, Modeling, and Evaluation.

Submit .ipynb file along with a document that consists the explanation and justification of the results.
Note: No Plagiarism