Develop a python code for analysis of tweeter dataset based on keywords and classify them
Budget: ₹1,500 – ₹12,500 INR
Labelled twitter dataset will be provide based on keyword the tweet have to be analysed and the tweet with the keywords is to the kept others have to removed and the result is to be saved as csv file.
1) preprocessed text corpus into uni-grams and bigram’s bag-of-words vector
2) use tf-idf (short for term frequency-inverse document frequency) and BM25 to calculate the weight for each term to obtain terms’ vector
3) classify the document as informative or not based on the results and label it as not_medical
4) classify them using transformer model
5) provide best accuracy (+90%)
6)provide confusion matrix, f1 score, precision, recall, false positive, true negative
7)Provide explanation of the ipython code In the notebook
after finalizing the budget no further reward will be provided and payment will be done only after completion of task
1) preprocessed text corpus into uni-grams and bigram’s bag-of-words vector
2) use tf-idf (short for term frequency-inverse document frequency) and BM25 to calculate the weight for each term to obtain terms’ vector
3) classify the document as informative or not based on the results and label it as not_medical
4) classify them using transformer model
5) provide best accuracy (+90%)
6)provide confusion matrix, f1 score, precision, recall, false positive, true negative
7)Provide explanation of the ipython code In the notebook
after finalizing the budget no further reward will be provided and payment will be done only after completion of task