Data collection and Refinement for Machine Learning Model
Budget: ₹1,500 – ₹12,500 INR
Project Overview:
We are looking for a skilled data analyst or data engineer to collect, clean, and refine external datasets to help us build a robust machine learning model for predicting student behavior in our social app. Our app, UNIGUYS, focuses on connecting students in the UK by helping them meet new friends, find study groups, and match based on social vibes.
The goal is to gather pre-fed data that will enhance our model’s ability to understand cultural norms, academic timelines, social behaviors, and other key factors that influence student behavior. This data will be crucial in predicting user preferences, moods, and interactions more accurately.
If you have expertise in gathering, processing, and refining large-scale datasets, we’d love to hear from you!
Key Responsibilities:
• Collect and aggregate external datasets relevant to the behavior of students, including but not limited to:
• Cultural Norms and Communication Styles: Data on communication preferences based on nationality and cultural background.
• Academic Calendars: Schedules from UK universities, including exam periods, breaks, and major academic events.
• Social Trends and Interests: Data on popular student interests, hobbies, and trends.
• Psychological and Behavioral Data: Personality traits (Big Five, MBTI), stressors, and behavior patterns from public datasets or research.
• Third-Party Data Sources: Integration of APIs from platforms like Spotify, YouTube, or Coursera for user interest prediction.
• Clean and preprocess the data, ensuring accuracy, removing duplicates, and handling missing data points.
• Structure the data into usable formats such as CSV files, SQL databases, or other required formats to integrate into our ML model.
• Document the data collection process, including source references, data extraction methods, and any potential limitations.
Requirements:
• Proven experience in data collection, cleaning, and preprocessing.
• Familiarity with academic, social, and cultural data sources (government databases, research papers, public APIs).
• Experience with data extraction tools and techniques (e.g., web scraping, API integration).
• Strong skills in data analysis and structuring using Python (Pandas, NumPy), SQL, or Excel.
• Knowledge of ethical data sourcing and compliance with privacy regulations (e.g., GDPR).
• Excellent documentation skills to ensure data traceability and reproducibility.
Preferred Qualifications:
• Familiarity with behavioral data, student trends, or working on educational/social platforms.
• Experience in collecting large datasets from public sources, APIs, or scraping academic resources.
• Experience integrating data into machine learning pipelines is a plus.
Deliverables:
• Cleaned, structured datasets ready for use in our machine learning model.
• Detailed documentation on data sources, collection methods, and preprocessing steps.
• Ongoing communication with our development team to ensure data quality and alignment with project goals.
We are looking for a skilled data analyst or data engineer to collect, clean, and refine external datasets to help us build a robust machine learning model for predicting student behavior in our social app. Our app, UNIGUYS, focuses on connecting students in the UK by helping them meet new friends, find study groups, and match based on social vibes.
The goal is to gather pre-fed data that will enhance our model’s ability to understand cultural norms, academic timelines, social behaviors, and other key factors that influence student behavior. This data will be crucial in predicting user preferences, moods, and interactions more accurately.
If you have expertise in gathering, processing, and refining large-scale datasets, we’d love to hear from you!
Key Responsibilities:
• Collect and aggregate external datasets relevant to the behavior of students, including but not limited to:
• Cultural Norms and Communication Styles: Data on communication preferences based on nationality and cultural background.
• Academic Calendars: Schedules from UK universities, including exam periods, breaks, and major academic events.
• Social Trends and Interests: Data on popular student interests, hobbies, and trends.
• Psychological and Behavioral Data: Personality traits (Big Five, MBTI), stressors, and behavior patterns from public datasets or research.
• Third-Party Data Sources: Integration of APIs from platforms like Spotify, YouTube, or Coursera for user interest prediction.
• Clean and preprocess the data, ensuring accuracy, removing duplicates, and handling missing data points.
• Structure the data into usable formats such as CSV files, SQL databases, or other required formats to integrate into our ML model.
• Document the data collection process, including source references, data extraction methods, and any potential limitations.
Requirements:
• Proven experience in data collection, cleaning, and preprocessing.
• Familiarity with academic, social, and cultural data sources (government databases, research papers, public APIs).
• Experience with data extraction tools and techniques (e.g., web scraping, API integration).
• Strong skills in data analysis and structuring using Python (Pandas, NumPy), SQL, or Excel.
• Knowledge of ethical data sourcing and compliance with privacy regulations (e.g., GDPR).
• Excellent documentation skills to ensure data traceability and reproducibility.
Preferred Qualifications:
• Familiarity with behavioral data, student trends, or working on educational/social platforms.
• Experience in collecting large datasets from public sources, APIs, or scraping academic resources.
• Experience integrating data into machine learning pipelines is a plus.
Deliverables:
• Cleaned, structured datasets ready for use in our machine learning model.
• Detailed documentation on data sources, collection methods, and preprocessing steps.
• Ongoing communication with our development team to ensure data quality and alignment with project goals.