Program and database for uploading and analyzing Chinese media sources -- 2

Job ID: 33103976

Budget: $1,500 – $3,000 USD

The database for collecting articles from the Chinese media and an algorithm (program) for the subsequent processing of the downloaded materials:
Functions:
1. The database must collect data (news/articles) from Chinese media: 新华网,人民网,咸宁新闻网,上观,新闻中心等等
2. The algorithm (program) must process found information based on 7 criterions/guidelines: source of information, date of publication, headline of news/article, character of publication (negative, positive, neutral), keywords, text, link to this news/article.
3. The algorithm (program) must identify the character of the information analyzing the headline of the article.

For instance, if the headline includes such words as invasion (侵入) or occupation (占领), then the character of the article is negative or “-1”. If the headline includes such words as “military operation” (军事行动), then the article has neutral character or “o”. If the article has such words as “military operation for democracy” (争取民主的军事行动), then the article has positive character or “+1”.
4. The program must renew the database or upload new materials by a request from Chinese media.
5. The database must contain news reports from Chinese media in the period of 2016-2021.
6. The interface of the algorithm (program) must be in English, whereas uploaded and processed data can be in Chinese.
7. It is very important that the database and the program (algorithm) should work outside China, i.e. users have an opportunity to use the algorithm (program) in Russia or in other states.
An example of similar database is the GDELT Project. The GDELT Project monitors a news report breaking anywhere the world. The GDELT Database contains the information processed from media sources.