Microbiome-Depression ML Data Analysis

Job ID: 39914063

Budget: $250 – $750 USD

I want to hand over a curated, de-identified microbiome dataset that we generated at Evolve Genomix with qPCR and next-generation sequencing. Every sample is paired with rich clinical metadata and repeated depression-symptom scores.

The task is to run a full machine-learning analysis service: clean and normalize abundance tables, build predictive or association models that link microbial taxa to current and future depression status, and document the entire pipeline so our in-house team can reproduce and extend it later. Python (scikit-learn, XGBoost, PyTorch, or similar) or R (caret, tidymodels, randomForest) are all acceptable as long as the code is well commented and version-controlled.

Key deliverables
• Pre-processing scripts that transform raw count data into model-ready features (e.g., CLR, rlog, or other compositional methods)
• At least one robust supervised model with cross-validation, performance metrics, and an explanation of feature importance or SHAP values
• Visualizations that make findings intuitive for clinicians (ROC curves, taxa effect plots, dimensionality reductions)
• A short technical report summarizing methods, results, and biological interpretation



Timeline and meeting cadence are flexible; milestones can be arranged around data exploration, model development, and final reporting. If you have a track record applying ML to microbiome or similar omics data, I’d like to see examples of past analyses or preprints when we discuss next steps.