Excel assignment -- 2
Budget: $10 – $30 USD
I am looking for an Excel expert to help me with an advanced data analysis project. The main purpose of the project is data analysis and I do not have any specific preference for Excel function or tool to be used. The analysis is expected to be at an advanced level of complexity. The ideal freelancer for this project should have experience in advanced Excel data analysis techniques and be proficient in using Pivot Tables and VLOOKUP functions. Attention to detail and strong analytical skills are a must.
This is the question text :
employees provided when they originally applied for a job at the firm. For each employee, the following variables are listed: Promoted (1 if promoted within 10 years, 0 otherwise), GPA (college GPA at graduation), Sports (number of athletic activities during college), and Leadership (number of leadership roles in student organizations).
Create a random forest ensemble classification tree model. Select two predictor variables randomly to construct each weak learner. What is the overall accuracy rate of the model on the validation data? What is the AUC value of the model? Which is the most important predictor variable? Score the new cases in the HR_Score worksheet using the random forest ensemble classification tree model. How many new employees in the data set will likely be promoted within 10 years based on a cutoff probability value of 0.5?
Question 5.
In recent years, medical research has incorporated the use of data analytics to find new ways to detect heart disease in its early stage. Medical doctors are particularly interested in accurately identifying high-risk patients so that preventive care and intervention can be administered in a timely manner. The accompanying data file shows a patient’s age (Age), blood pressure (BP Systolic and BP Diastolic), and BMI, along with an indicator of whether or not the patient has heart disease (Disease = 1 if heart disease, 0 otherwise).
Create a bagging ensemble classification tree model to predict whether a patient has heart disease. What is the overall accuracy rate on the validation data? Create a boosting ensemble classification tree model. What is the overall accuracy rate of the model on the validation data? Compare the two ensemble models. Which model shows more robust performance according to the AUC value?
And mu budget around 15 usd thats all
This is the question text :
employees provided when they originally applied for a job at the firm. For each employee, the following variables are listed: Promoted (1 if promoted within 10 years, 0 otherwise), GPA (college GPA at graduation), Sports (number of athletic activities during college), and Leadership (number of leadership roles in student organizations).
Create a random forest ensemble classification tree model. Select two predictor variables randomly to construct each weak learner. What is the overall accuracy rate of the model on the validation data? What is the AUC value of the model? Which is the most important predictor variable? Score the new cases in the HR_Score worksheet using the random forest ensemble classification tree model. How many new employees in the data set will likely be promoted within 10 years based on a cutoff probability value of 0.5?
Question 5.
In recent years, medical research has incorporated the use of data analytics to find new ways to detect heart disease in its early stage. Medical doctors are particularly interested in accurately identifying high-risk patients so that preventive care and intervention can be administered in a timely manner. The accompanying data file shows a patient’s age (Age), blood pressure (BP Systolic and BP Diastolic), and BMI, along with an indicator of whether or not the patient has heart disease (Disease = 1 if heart disease, 0 otherwise).
Create a bagging ensemble classification tree model to predict whether a patient has heart disease. What is the overall accuracy rate on the validation data? Create a boosting ensemble classification tree model. What is the overall accuracy rate of the model on the validation data? Compare the two ensemble models. Which model shows more robust performance according to the AUC value?
And mu budget around 15 usd thats all