Python Algorithm for Data Error Identification
Budget: $10 – $30 USD
I need a python algorithm to identify errors in large set of data with I will provide upon successful bid.
I am implementing a subsidised school-lunch and parental support programme
to promote learning outcomes of children, targeting particularly Grade 1 to Grade 3 students in
Tanzania. The programme is currently being tested through an impact evaluation involving a group
of 30 participant schools, as well as a set of 30 control schools that will receive the program at a
later stage. The endline evaluation involves collecting data from schools over a one-month
period, focusing on measuring children’s learning as proxied by the Early Grade Reading and
Mathematics Assessments tools.
The data collection has recently kicked off, and a scheduled break at the start has been included
in the data collection process to allow for any necessary data corrections.
In your assignment, you have received a dataset coming from the first two days of data collection.
Your tasks are as follows:
1. Identifying follow-up actions before the data collection resumes: The data collection
is set to resume in 4 days’ time. Your first task is to check the incoming data and survey
tool programmed by a coding team and catch any notable errors that may compromise
data quality going forward.
2. Provide preliminary analysis on attrition rates and sample balance: Your second task
is to prepare a table showing comparing mean age, Mathematics and English test scores,
gender ratios, and sample sizes of students in the Treatment and Control samples
respectively.
Datasets:
- "student_scores.csv" – this dataset includes the raw export of materials collected through the CAPI app
- "school_sample.csv" – the school sample information, including a list of treatment and control schools
Survey Tool reference:
- "Student Survey_Endline.docx" – the paper questionnaire used in the survey suggested and shared by a Research Consultant
- " student_scores_CAPI_printable.pdf" – the CAPI tool logic used for collecting the data of the “student_scores.csv” dataset
I am implementing a subsidised school-lunch and parental support programme
to promote learning outcomes of children, targeting particularly Grade 1 to Grade 3 students in
Tanzania. The programme is currently being tested through an impact evaluation involving a group
of 30 participant schools, as well as a set of 30 control schools that will receive the program at a
later stage. The endline evaluation involves collecting data from schools over a one-month
period, focusing on measuring children’s learning as proxied by the Early Grade Reading and
Mathematics Assessments tools.
The data collection has recently kicked off, and a scheduled break at the start has been included
in the data collection process to allow for any necessary data corrections.
In your assignment, you have received a dataset coming from the first two days of data collection.
Your tasks are as follows:
1. Identifying follow-up actions before the data collection resumes: The data collection
is set to resume in 4 days’ time. Your first task is to check the incoming data and survey
tool programmed by a coding team and catch any notable errors that may compromise
data quality going forward.
2. Provide preliminary analysis on attrition rates and sample balance: Your second task
is to prepare a table showing comparing mean age, Mathematics and English test scores,
gender ratios, and sample sizes of students in the Treatment and Control samples
respectively.
Datasets:
- "student_scores.csv" – this dataset includes the raw export of materials collected through the CAPI app
- "school_sample.csv" – the school sample information, including a list of treatment and control schools
Survey Tool reference:
- "Student Survey_Endline.docx" – the paper questionnaire used in the survey suggested and shared by a Research Consultant
- " student_scores_CAPI_printable.pdf" – the CAPI tool logic used for collecting the data of the “student_scores.csv” dataset