Indian College Data Scraping

Job ID: 39733132

Budget: ₹100 – ₹400 INR

I need a clean, up-to-date dataset covering every college in India that can be discovered on public-facing web sources such as AICTE, UGC, NIRF, individual university lists, and the institutions’ own sites.

What I expect in the final hand-off:
• A well-structured file (CSV is ideal, but please also provide JSON so I can feed it straight into an API).
• Core fields for each college—name, full postal address, state, district, phone, email, website, and affiliation / accreditation tags—plus a concise list of the main academic programs offered.
• No duplicates, no placeholder rows, and every record aligned to the same column order.

You are free to choose the stack (Python + BeautifulSoup/Scrapy, Selenium, Playwright, etc.). Just ship the scripts or notebooks alongside brief run instructions so I can reproduce or refresh the crawl later.

Accuracy matters more than sheer volume, so build in validation checks, de-duplication and sensible rate-limiting to stay within site terms. Flag any sites that block scraping so we can decide on alternatives. Once the first 1,000 rows look solid, we’ll green-light the complete pull.

If you enjoy data wrangling and have tackled Indian education portals before, this should be straightforward. I’m ready to review a quick sample as soon as you have one.