Build a large dataset of 15,000+ Indian schools & colleges (name, contact, courses, fees, etc.)
Budget: ₹600 – ₹1,500 INR
I need someone to collect, clean and compile a comprehensive list of Indian educational institutions (schools & colleges) — about 15,000+ entries — with detailed info: name, website, phone, email, logo/photo link, about/description, founded year, courses offered, academic fees (if available), approval/affiliation (AICTE, BCI, PCI, board types, etc.), facilities, hostel fee (if any), institution type, admission criteria/eligibility, social media links, ranking/awards (if publicly available).
What I need from you:
- Find reliable, publicly available sources (institution sites, government/board listings, public directories).
- Collect the data and compile into a structured format (CSV/Excel/JSON) with clean columns.
- Validate and clean data (avoid duplicates, mark missing/unknown fields clearly).
- Document your data sources and date of extraction, so transparency is maintained.
Required skills:
- Web research/data gathering/scraping (if legally allowable) or manual collection.
- Data cleaning/structuring/normalization.
- Basic CSV / spreadsheet/database skills.
Deliverable:
A dataset file with 15,000+ rows and columns for all required fields (or as many as possible), plus a “source log” listing where each entry was collected from.
Important note:
Only use publicly available information. Avoid copying copyrighted descriptive content verbatim (e.g. long descriptions, images, logo files with restricted license).
What I need from you:
- Find reliable, publicly available sources (institution sites, government/board listings, public directories).
- Collect the data and compile into a structured format (CSV/Excel/JSON) with clean columns.
- Validate and clean data (avoid duplicates, mark missing/unknown fields clearly).
- Document your data sources and date of extraction, so transparency is maintained.
Required skills:
- Web research/data gathering/scraping (if legally allowable) or manual collection.
- Data cleaning/structuring/normalization.
- Basic CSV / spreadsheet/database skills.
Deliverable:
A dataset file with 15,000+ rows and columns for all required fields (or as many as possible), plus a “source log” listing where each entry was collected from.
Important note:
Only use publicly available information. Avoid copying copyrighted descriptive content verbatim (e.g. long descriptions, images, logo files with restricted license).
Related categories:
Data Entry
Excel
Web Scraping
Web Search
JSON
Data Scraping
Data Collection
Data Management