Optimal Clustering for Text Dataset

Job ID: 38963540

Budget: £10 – £20 GBP

I'm seeking an expert who can determine the ideal number of clusters for a text dataset. The project involves k-means clustering using k-fold cross-validation, complete with preprocessing and evaluation in python only.

Key Tasks:
- Use the k-means algorithm exclusively
- Conduct k-fold cross-validation
- Preprocess the dataset through tokenization, stop words removal, and lemmatization
- Evaluate clustering results using the silhouette score and elbow method

I'm open to various origins for the dataset, such as customer reviews, social media posts, or scientific articles. The key is your ability to deliver the necessary code and provide a clear explanation of the results.

Ideal Skills and Experience:
- Proficient in Python or R
- Extensive experience with k-means clustering
- Strong background in text data preprocessing and analysis
- Familiarity with k-fold cross-validation
- Ability to interpret and explain clustering results