Project on Bigdata

Job ID: 32311400

Budget: $10 – $30 USD

The course is organized around big data technologies and modeling pipelines, and the project also
focuses in these areas. The students can choose from one of the following options:
o Dataiku implementation of Data Science Pipeline OR
o Develop a data preparation and modeling pipeline in language of their choice (either Python
or R).
Knowledge of these languages is not a prerequisite of the course and was also not discussed.
Hence if you choose this option, I expect that you have background with these languages and
will be able to work on them. Please choose this option ONLY if you already know these
languages. I will be happy to assess and provide feedback.
• When using either of these options students are expected to meet the requirements laid in the
rubric.
• The project aims that a student will work on
o collecting or accessing or retrieving sample big dataset
o preparing data for analysis
o identifying the type of statistical and analytical models to apply
o presenting the results by packaging this in a `story-telling’ project report
• The project is worth 100 points (18% of the course grade)

Submission
• You can explore these webpages for datasets:
o UC Irvine Machine Learning Repository - https://archive.ics.uci.edu/ml/index.php
o Kaggle Datasets - https://www.kaggle.com/datasets
o You are not restricted to these datasets.
o These datasets are public, and more than one student may end up using the same dataset.
Please make sure that you submit your original implementation. In case of any dispute
Academic dishonesty policy will be applicable.
o Your dataset should have at least 10000 observations.
• The submission should include:
o Dataiku Project file (Exported from Project Homepage) OR Python/R codes with datasets (if
you choose this option)
o Final report describing your questions, methods and results

The students will choose a dataset (some resources are provided in the last section of this handout) and
conduct the following analyses:
• List ONE research question that could be a supervised or unsupervised learning model
(regression, classification, clustering, anomaly detection etc.)
• Conduct ONE data preparation step such as data type transformation, join, split, sort etc.
• Implement you chosen model.
• A PDF generated from a typed Word document of 3-5 pages
o Describe your data
o Describe your research question
o Discuss the data preparation
o The results of the model
o Report should be well presented describing your data analytics task clearly