Data Sciences project

Job ID: 37428336

Budget: $30 – $250 USD

1. Detailed and not summary (aggregated) dataset. The dataset must also have high number of rows.
2. Aligned with Saudi National Priorities Topics (see attached). In your submission and presentations, you must specify the topic. For example: Economies of the Future -> Artificial intelligence -> Dialect language detection and adaptation (especially for semantically complex languages, such as Arabic and Mandarin).

In the project, you will be doing the following:
1. For requirement #1, you have two choices:

Link for dataset:
* https://research.google/tools/datasets/
* https://paperswithcode.com/datasets
* https://archive.ics.uci.edu/
* https://msropendata.com/

* Use an existing dataset and manually label it. Here, you can select any high-qualitydataset. Therefore, you cannot select a dataset from Kaggle uploaded by an unknown person. You can only select datasets collected by organizations or named individuals working for well-known organizations. For labeling, it must be one that is essential for the problem you are targeting. Generally, you should be looking for adding labels to text or vision data. The labeling must be done by all team members with details for all labeling procedures and steps followed to ensure that your labels are of high quality. Please contact me to confirm that your labels are suitable for the problem. Do not start labeling without consulting me first.
1. Only if needed, detail any cleaning or processing done to prepare the dataset for subsequent phases. Do not clean columns that you will not use. Also, detail any potential data quality issues or mistakes.
2. Using Tableau, create summary statistics and data visualizations that will help show the dataset and various descriptive information. You can create individual charts, a dashboard, or a story (or all) ,Share the link for tableau and share tableau file.
3. Write at least two interesting hypotheses and then test them using the appropriate tests. Specify why they interest you and provide background information that will help us understand the problem in general. If you do 1.b, at least one of the hypotheses must have the labels you created.
4. Use an existing model trained by others on your dataset. Visit hugginface and paperswithcode and look for models that you can try and test. Then, using your own data, evaluate how the mode performs on your own subset. If you want to also develop your own models or tune existing ones to improve results you can do that (this part is optional).
Submissions:

Presentation powerpoint:

Present the entire project. In the submission, you can submit your work as a powerpoint presentation submit and share your code. If done using colab (this is preferred), export your code as a pdf and also share it and include the URL in the submission
Related categories: Python Powerpoint Report Writing Data Science