Thesis with R Statistical Analysis
Budget: $250 – $750 USD
The study is ready to move from raw data to a full, publish-ready manuscript. I will supply the dataset and preliminary research questions; your task is to transform them into a 40-plus-page thesis that we can confidently submit to an academic committee.
Scope of work
• Data preparation: clean and document the dataset in R (tidyverse/dplyr preferred), retaining a reproducible log of every transformation.
• Analysis: run descriptive, inferential, and cluster techniques as appropriate—factor analysis, k-means or hierarchical clustering, regression models, and any post-hoc tests needed. All code must be annotated and delivered as .R or .Rmd files that knit without errors.
• Visualisation: use ggplot2 (or comparable libraries) for clear graphs, advanced cluster plots, heatmaps, and any additional plots that clarify findings. Each figure needs an explanatory caption ready for the thesis.
• Literature review: 25–35 recent peer-reviewed sources synthesised in either APA or Harvard style (I will confirm which one before you begin).
• Write-up: introduction, methodology, results, discussion, limitations, and conclusion, formatted in Word and exported to PDF. The narrative should flow logically, interpret statistics in plain language, and tie back to the literature.
• Reference list: auto-generated from your citation manager so all in-text citations are accounted for.
Acceptance criteria
1. R scripts reproduce every table and figure on my machine.
2. Thesis document exceeds 40 pages (excluding appendices) with coherent academic structure.
3. Referencing style is consistent and error-free.
4. All visuals are high-resolution and embedded in the report with captions.
Send a brief outline of your intended analytical approach and timeline so we can confirm milestones and get started.
Scope of work
• Data preparation: clean and document the dataset in R (tidyverse/dplyr preferred), retaining a reproducible log of every transformation.
• Analysis: run descriptive, inferential, and cluster techniques as appropriate—factor analysis, k-means or hierarchical clustering, regression models, and any post-hoc tests needed. All code must be annotated and delivered as .R or .Rmd files that knit without errors.
• Visualisation: use ggplot2 (or comparable libraries) for clear graphs, advanced cluster plots, heatmaps, and any additional plots that clarify findings. Each figure needs an explanatory caption ready for the thesis.
• Literature review: 25–35 recent peer-reviewed sources synthesised in either APA or Harvard style (I will confirm which one before you begin).
• Write-up: introduction, methodology, results, discussion, limitations, and conclusion, formatted in Word and exported to PDF. The narrative should flow logically, interpret statistics in plain language, and tie back to the literature.
• Reference list: auto-generated from your citation manager so all in-text citations are accounted for.
Acceptance criteria
1. R scripts reproduce every table and figure on my machine.
2. Thesis document exceeds 40 pages (excluding appendices) with coherent academic structure.
3. Referencing style is consistent and error-free.
4. All visuals are high-resolution and embedded in the report with captions.
Send a brief outline of your intended analytical approach and timeline so we can confirm milestones and get started.