Machine learning Project
Budget: €250 – €750 EUR
Extraction and keywording of texts highly relevant to social media discourses on COVID-19 and conspiracy myths.
In the Digital Hate research project, we examine social media discourses in the context of conspiracy narratives about COVID-19. Due to their brevity, social media posts such as tweets often contain little direct information and instead link to websites such as blogs, newspaper portals, or video platforms.
In this study project, the content of frequently linked news websites will be extracted and tagged by keywords. Specifically, this requires the following work steps:
1. extraction of URLs from social media posts (these are already in a database) and finding the most frequent ones that are not typical video hosting platforms. We want to focus on news sites.
2. application of machine learning algorithms like Web2Text to remove boilerplate code and extract the main content. (Blog Post on Web2Text with code samples: https://xaviergeerinck.com/post/ai/web2text)
3. automatic extraction of keywords from the cleaned texts using Natural Language Processing and Machine Learning Algorithms (will be specified later in the project).
4. optional: Flag the URLs using an external list of typical fake news websites (this will require additional data from external sources such as the organization Newsguard but it is not clear whether we can access those).
The project requires an interest in data processing and programming with Python. One difficulty is that most of the texts are in German, however there is support in the project for this.
NOTE: use juypter notebook platform and use python coding, also need project report at least 25 pages . please find the attached Datasheet
In the Digital Hate research project, we examine social media discourses in the context of conspiracy narratives about COVID-19. Due to their brevity, social media posts such as tweets often contain little direct information and instead link to websites such as blogs, newspaper portals, or video platforms.
In this study project, the content of frequently linked news websites will be extracted and tagged by keywords. Specifically, this requires the following work steps:
1. extraction of URLs from social media posts (these are already in a database) and finding the most frequent ones that are not typical video hosting platforms. We want to focus on news sites.
2. application of machine learning algorithms like Web2Text to remove boilerplate code and extract the main content. (Blog Post on Web2Text with code samples: https://xaviergeerinck.com/post/ai/web2text)
3. automatic extraction of keywords from the cleaned texts using Natural Language Processing and Machine Learning Algorithms (will be specified later in the project).
4. optional: Flag the URLs using an external list of typical fake news websites (this will require additional data from external sources such as the organization Newsguard but it is not clear whether we can access those).
The project requires an interest in data processing and programming with Python. One difficulty is that most of the texts are in German, however there is support in the project for this.
NOTE: use juypter notebook platform and use python coding, also need project report at least 25 pages . please find the attached Datasheet