Fraud Detection Binary Classifier
Budget: ₹1,000 – ₹2,000 INR
I have a structured spreadsheet dataset (10-50 columns) where each row is tagged as either fraudulent or legitimate, and I want to turn it into a reliable fraud-detection pipeline. Your task begins with a quick exploratory analysis to spot class imbalance, outliers, and any feature quirks, then moves on to feature engineering, model selection, training, and evaluation. I’m comfortable with tried-and-true Python tooling—pandas for wrangling, scikit-learn or XGBoost for modelling, and seaborn/Matplotlib for visual insight—but I’m open to your preferred libraries if they achieve better performance or transparency.
Deliverables
• Notebook or Python scripts that load the raw CSV files, clean and prep the data, train the binary classifier, and output predictions
• A concise report (PDF or Markdown) explaining methodology, performance metrics (AUC-ROC, precision-recall, confusion matrix), and any recommendations for deployment or further tuning
• Reproducibility: a requirements.txt or environment.yml and clear run instructions
Acceptance criteria
We can discuss, need higher precision and recall.
Deliverables
• Notebook or Python scripts that load the raw CSV files, clean and prep the data, train the binary classifier, and output predictions
• A concise report (PDF or Markdown) explaining methodology, performance metrics (AUC-ROC, precision-recall, confusion matrix), and any recommendations for deployment or further tuning
• Reproducibility: a requirements.txt or environment.yml and clear run instructions
Acceptance criteria
We can discuss, need higher precision and recall.
Related categories:
Python
Software Architecture
Machine Learning (ML)
Data Mining
Data Science
Data Visualization
Data Analysis
Pandas