Visualizing Chinese Text: Wordcloud & Topic modeling
Budget: $30 – $250 USD
We are looking for someone specialized in NLP in the Chinese language. We would like to visualize a text corpus using (1) wordcloud and (2) topic modeling methods (e.g. LDA). The corpus is stored in the “text” column in a csv file with each row corresponding to a post/document. See the sample text in the attached fig1. The texts are mainly in Simplified Chinese but may occasionally include English words.
You need to separately pre-process Chinese and English texts by
1. Remove urls
2. Use the Jieba package in R or Python to segment texts in Simplified Chinese.
3. Remove stopwords and non-words (e.g. punctuation, numbers etc.)
4. Word stemming (mostly applying to English words)
You must be able to complete the task and provide the script within 1-2 days using R or Python. Please do not contact me if you cannot finish the task within that time frame. The desired outputs should resemble the attached fig2 (word cloud) & fig3 (topic detection).
Please provide in your message an overview of how you plan to approach the task. Thanks!
You need to separately pre-process Chinese and English texts by
1. Remove urls
2. Use the Jieba package in R or Python to segment texts in Simplified Chinese.
3. Remove stopwords and non-words (e.g. punctuation, numbers etc.)
4. Word stemming (mostly applying to English words)
You must be able to complete the task and provide the script within 1-2 days using R or Python. Please do not contact me if you cannot finish the task within that time frame. The desired outputs should resemble the attached fig2 (word cloud) & fig3 (topic detection).
Please provide in your message an overview of how you plan to approach the task. Thanks!