Node.js Scholarship Data Scraper
Budget: ₹1,500 – ₹12,500 INR
I need a small, reliable Node.js script that automatically pulls fresh scholarship information from government sites, university pages, and the main private-scholarship portals I follow. The job is strictly about collecting data; I’m not looking for automated filtering or application submission at this stage.
Here’s what I expect:
• Crawl each source on a repeatable schedule, discover new or updated listings, and capture core fields such as scholarship name, provider, eligibility summary, award amount, key dates, and the original URL.
• Insert the results directly into my database (I can provision PostgreSQL or MongoDB—whichever you’re more comfortable with).
• Flag duplicates to prevent double-entry and log any pages that fail to parse so I can review them later.
• Keep the code clean, commented, and set up so I can add new sources easily.
If you’ve worked with cheerio, puppeteer, or similar Node scraping tools, this should be straightforward. Once the initial crawler is stable, there’s room to expand into filtering and auto-apply features, so clean architecture matters.
Let me know the stack you propose, any rate limits we should watch for, and how soon you can deliver a first working version.
Here’s what I expect:
• Crawl each source on a repeatable schedule, discover new or updated listings, and capture core fields such as scholarship name, provider, eligibility summary, award amount, key dates, and the original URL.
• Insert the results directly into my database (I can provision PostgreSQL or MongoDB—whichever you’re more comfortable with).
• Flag duplicates to prevent double-entry and log any pages that fail to parse so I can review them later.
• Keep the code clean, commented, and set up so I can add new sources easily.
If you’ve worked with cheerio, puppeteer, or similar Node scraping tools, this should be straightforward. Once the initial crawler is stable, there’s room to expand into filtering and auto-apply features, so clean architecture matters.
Let me know the stack you propose, any rate limits we should watch for, and how soon you can deliver a first working version.
Related categories:
JavaScript
Web Scraping
NoSQL Couch & Mongo
Node.js
PostgreSQL
MongoDB
Data Collection
Data Management