Enhanced Insider Threat Detection Through Preprocessing

Job ID: 37754320

Budget: $30 – $250 USD

My project brings a focus on developing an advanced insider threat detection model several objectives:
- Proprocess user activity logs (Cleaning & Merging)
- Feature Extraction (Convert categorical features to word embedding)
- Train sequential models (mainly transformer models) using state-of-the-art method from several papers

For this project, I'll be utilizing user activity logs from CMU CERT dataset (r4.2) as my main data source, a plentiful and rich trove of metadata to draw from for pattern identification.

The ideal candidate for this project will have experience with:
- Insider threat detection
- Machine Learning frameworks (Pytorch, Pandas, Numpy, Hugging Face)
- Implementing codes from other state-of-the-art papers on custom dataset
- Handling and manipulating user activity logs

The centerpiece method for this project will be Feature Extraction. Therefore, I am looking for someone who has substantial experience in this approach to data preprocessing. This should include knowing how to identify and extract meaningful features from the user activity logs to feed into our threat detection model.
Related categories: Machine Learning (ML) Pytorch Pandas Hugging Face