Production Event Scraper Development
Budget: $250 – $750 USD
PROJECT: ForaHub - Web Platform Finalization, Production Scraper System, AI Integration, and Mobile App Readiness
ABOUT THE PRODUCT
ForaHub (forahub.org) is a global directory and discovery platform for development, humanitarian, climate, health, policy, and academic events. Its audience is institutional: UN agencies, NGOs, foundations, universities, research institutions, development partners, and government bodies. Users discover events through a searchable directory and an interactive map. Organizations can claim and manage their own profiles and events.
The platform is already built and functional, using Next.js (App Router), Supabase (Postgres, Auth, Storage, Row Level Security), Vercel, Tailwind CSS, and a Leaflet map. It has recently undergone a security, performance, and accessibility hardening pass. This is a one-month contract to take the product from "functional" to "fully launch-ready and production-grade."
PRIMARY DELIVERABLE: A SERIOUS, PRODUCTION-GRADE SCRAPER SYSTEM
This is the single most important part of the engagement. ForaHub's value depends on a continuously growing, accurate, well-structured directory of events drawn automatically from many sources. The current scraper is not sufficient. We need someone who can architect and build a robust, scalable, maintainable scraping SYSTEM, not a simple script or a quick fix.
Scraper requirements:
- Collect a large, continuous volume of high-quality events from many diverse sources: organization websites, event platforms, calendars, JavaScript-rendered pages, PDFs and documents, RSS/iCal feeds, and APIs.
- Handle varied and messy source formats robustly, with a clean, extensible way to add new sources over time.
- Accurate field extraction: title, start and end dates, location, organizer, registration link, and description.
- Reliable deduplication across sources, geocoding of locations, and AI-assisted categorization and SDG tagging that is validated for accuracy.
- Reliability and scale: scheduled and automated runs, respectful crawling with rate-limiting, retries and graceful failure handling, resilience to source layout changes, and the ability to scale to thousands of events without degrading the site.
- Observability: logging, monitoring, run reports, clear error surfacing, and alerting or a status view for failed sources.
- Operate within a modest, predictable monthly AI/processing/compute budget.
- Integrate cleanly with the existing events and organizations data model with no duplicates or conflicts.
- Documented architecture, sources, and data flow, with clear instructions for adding, debugging, and maintaining sources after handover.
Proven, demonstrable scraping and data-pipeline experience is mandatory. Applicants must show real examples of scraping systems they have built, including the sources handled, how they dealt with anti-bot measures and JavaScript-rendered content, deduplication, and reliability. Generic claims without specifics will not be considered.
OTHER DELIVERABLES
1. Web platform finalization
- Complete remaining or partially built features: organization management, claim and verification flows, team accounts, auto-publish, recurring events, and analytics.
- Fix outstanding bugs and edge cases. Ensure all critical user journeys work end to end: sign-up, organization claim, co-manager invitation and acceptance, event submission and publishing, recurring series, and account deletion.
- Cross-browser testing. Clear, user-friendly error and empty states.
2. AI integration (finalize and apply guardrails)
- Finalize and complete the AI-driven features so they are production-ready: AI parsing and extraction in the scraper, AI-based event categorization and SDG tagging, and any AI assistant features in the product.
- Ensure each AI integration is fully wired end to end, accurate, and reliable, not left half-built or experimental.
- Implement AI responsibly: cost-controlled within a predictable monthly budget, with sensible fallbacks when the AI service is unavailable or returns poor output.
- AI must never fabricate data or display misleading or fake activity. AI outputs that affect users must be validated for accuracy.
- Document how each AI integration works, what it costs, and how to maintain it.
3. Mobile applications - app-store readiness
- Prepare the mobile application so it is ready to submit and deploy to the Apple App Store (iOS) and Google Play Store (Android): correct build configuration, app icons, splash screens, store metadata, permissions, and compliance with each store's submission requirements.
- Confirm consistent, stable behavior on both iOS and Android, and provide the builds and documentation needed for submission.
- Note: the Apple and Google developer accounts and the final store submission are the Product Owner's responsibility. The Consultant delivers the apps in a ready-to-submit, store-compliant state.
4. UI/UX and look and feel
- Elevate the overall visual design to an institutional, trusted, premium standard suitable for UN agencies, NGOs, foundations, and universities.
- Improve visual hierarchy, spacing, typography, consistency, and component styling across the platform.
- Strengthen empty states, loading states, and onboarding so the product never feels unfinished.
- Deliver a consistent, well-implemented light and dark mode across every page.
5. Performance, accessibility, and quality
- Maintain and improve performance: efficient queries, pagination, optimized assets, fast load times.
- Maintain and improve accessibility toward WCAG AA, which matters for the institutional audience.
- No regressions to existing security, Row Level Security, authentication, or data integrity. Permission checks must remain enforced server-side.
SUSTAINABILITY AND TESTING
- All work must be sustainable and built to last beyond this engagement: clean, modular, documented, and maintainable code, not short-term patches that create future debt.
- Testing is required for every deliverable: build cleanly with no new errors; manually verify all critical flows; test the scraper against real sources for output quality, deduplication, and categorization; test across major browsers and mobile screen sizes; verify light and dark mode; run performance and accessibility checks with no regression. Provide a short test summary with each milestone.
TECHNOLOGY STACK (must work within the existing stack, no unapproved rewrites)
Next.js (App Router) and React; Supabase (Postgres, Auth, Storage, Row Level Security); Tailwind CSS; Vercel; Leaflet. Web scraping, data pipelines, and AI/LLM API integration are required for the scraper and AI work.
REQUIRED SKILLS AND EXPERIENCE
years professional full-stack development experience, including at least 2 years building production web scrapers or data pipelines. Scraping experience is mandatory.
- Proven, demonstrable scraping ability with concrete examples (sources, anti-bot handling, JavaScript-rendered content, deduplication, reliability).
- Strong command of scraping and data-engineering tooling: headless browsers, HTML and structured-data parsing, scheduling and orchestration, queues, rate-limiting, retries, and monitoring.
- Strong skills in Next.js, React, Supabase/Postgres, Tailwind CSS, and Vercel.
- Responsible AI/LLM integration with cost control and fallbacks.
- UI/UX sensibility with a portfolio of clean, production-grade interfaces.
- Experience preparing iOS and Android apps for App Store and Play Store submission (strongly preferred).
- Fluent English, written and spoken, is mandatory. Clear, proactive communication is essential.
- A disciplined approach to testing and writing sustainable, maintainable, documented code.
ENGAGEMENT TERMS
- One-month, fixed-price contract.
- Paid in milestones: kickoff (after a paid week-one trial task), mid-point on demonstrated progress, and final on acceptance and merge of the completed work.
- A paid trial task in week one, focused on the scraper (for example, adding a new non-trivial source with correct extraction, deduplication, and error handling), is used to confirm fit before fuller commitment. The trial task is paid even if the engagement does not proceed.
- The Consultant must sign a Non-Disclosure Agreement, accept work-for-hire intellectual property terms (all work product owned by ForaHub), work on a separate Git branch, and submit work via pull request for review. No direct pushes or merges to the main branch.
- Access is granted progressively and on a least-privilege basis; production secrets are not shared. All access is revoked at the end of the engagement.
- No destructive actions (deleting data, dropping tables, rotating or exposing secrets, altering production environment variables) without written approval.
- Strong performance may lead to an ongoing or long-term arrangement, agreed separately.
ABOUT THE PRODUCT
ForaHub (forahub.org) is a global directory and discovery platform for development, humanitarian, climate, health, policy, and academic events. Its audience is institutional: UN agencies, NGOs, foundations, universities, research institutions, development partners, and government bodies. Users discover events through a searchable directory and an interactive map. Organizations can claim and manage their own profiles and events.
The platform is already built and functional, using Next.js (App Router), Supabase (Postgres, Auth, Storage, Row Level Security), Vercel, Tailwind CSS, and a Leaflet map. It has recently undergone a security, performance, and accessibility hardening pass. This is a one-month contract to take the product from "functional" to "fully launch-ready and production-grade."
PRIMARY DELIVERABLE: A SERIOUS, PRODUCTION-GRADE SCRAPER SYSTEM
This is the single most important part of the engagement. ForaHub's value depends on a continuously growing, accurate, well-structured directory of events drawn automatically from many sources. The current scraper is not sufficient. We need someone who can architect and build a robust, scalable, maintainable scraping SYSTEM, not a simple script or a quick fix.
Scraper requirements:
- Collect a large, continuous volume of high-quality events from many diverse sources: organization websites, event platforms, calendars, JavaScript-rendered pages, PDFs and documents, RSS/iCal feeds, and APIs.
- Handle varied and messy source formats robustly, with a clean, extensible way to add new sources over time.
- Accurate field extraction: title, start and end dates, location, organizer, registration link, and description.
- Reliable deduplication across sources, geocoding of locations, and AI-assisted categorization and SDG tagging that is validated for accuracy.
- Reliability and scale: scheduled and automated runs, respectful crawling with rate-limiting, retries and graceful failure handling, resilience to source layout changes, and the ability to scale to thousands of events without degrading the site.
- Observability: logging, monitoring, run reports, clear error surfacing, and alerting or a status view for failed sources.
- Operate within a modest, predictable monthly AI/processing/compute budget.
- Integrate cleanly with the existing events and organizations data model with no duplicates or conflicts.
- Documented architecture, sources, and data flow, with clear instructions for adding, debugging, and maintaining sources after handover.
Proven, demonstrable scraping and data-pipeline experience is mandatory. Applicants must show real examples of scraping systems they have built, including the sources handled, how they dealt with anti-bot measures and JavaScript-rendered content, deduplication, and reliability. Generic claims without specifics will not be considered.
OTHER DELIVERABLES
1. Web platform finalization
- Complete remaining or partially built features: organization management, claim and verification flows, team accounts, auto-publish, recurring events, and analytics.
- Fix outstanding bugs and edge cases. Ensure all critical user journeys work end to end: sign-up, organization claim, co-manager invitation and acceptance, event submission and publishing, recurring series, and account deletion.
- Cross-browser testing. Clear, user-friendly error and empty states.
2. AI integration (finalize and apply guardrails)
- Finalize and complete the AI-driven features so they are production-ready: AI parsing and extraction in the scraper, AI-based event categorization and SDG tagging, and any AI assistant features in the product.
- Ensure each AI integration is fully wired end to end, accurate, and reliable, not left half-built or experimental.
- Implement AI responsibly: cost-controlled within a predictable monthly budget, with sensible fallbacks when the AI service is unavailable or returns poor output.
- AI must never fabricate data or display misleading or fake activity. AI outputs that affect users must be validated for accuracy.
- Document how each AI integration works, what it costs, and how to maintain it.
3. Mobile applications - app-store readiness
- Prepare the mobile application so it is ready to submit and deploy to the Apple App Store (iOS) and Google Play Store (Android): correct build configuration, app icons, splash screens, store metadata, permissions, and compliance with each store's submission requirements.
- Confirm consistent, stable behavior on both iOS and Android, and provide the builds and documentation needed for submission.
- Note: the Apple and Google developer accounts and the final store submission are the Product Owner's responsibility. The Consultant delivers the apps in a ready-to-submit, store-compliant state.
4. UI/UX and look and feel
- Elevate the overall visual design to an institutional, trusted, premium standard suitable for UN agencies, NGOs, foundations, and universities.
- Improve visual hierarchy, spacing, typography, consistency, and component styling across the platform.
- Strengthen empty states, loading states, and onboarding so the product never feels unfinished.
- Deliver a consistent, well-implemented light and dark mode across every page.
5. Performance, accessibility, and quality
- Maintain and improve performance: efficient queries, pagination, optimized assets, fast load times.
- Maintain and improve accessibility toward WCAG AA, which matters for the institutional audience.
- No regressions to existing security, Row Level Security, authentication, or data integrity. Permission checks must remain enforced server-side.
SUSTAINABILITY AND TESTING
- All work must be sustainable and built to last beyond this engagement: clean, modular, documented, and maintainable code, not short-term patches that create future debt.
- Testing is required for every deliverable: build cleanly with no new errors; manually verify all critical flows; test the scraper against real sources for output quality, deduplication, and categorization; test across major browsers and mobile screen sizes; verify light and dark mode; run performance and accessibility checks with no regression. Provide a short test summary with each milestone.
TECHNOLOGY STACK (must work within the existing stack, no unapproved rewrites)
Next.js (App Router) and React; Supabase (Postgres, Auth, Storage, Row Level Security); Tailwind CSS; Vercel; Leaflet. Web scraping, data pipelines, and AI/LLM API integration are required for the scraper and AI work.
REQUIRED SKILLS AND EXPERIENCE
years professional full-stack development experience, including at least 2 years building production web scrapers or data pipelines. Scraping experience is mandatory.
- Proven, demonstrable scraping ability with concrete examples (sources, anti-bot handling, JavaScript-rendered content, deduplication, reliability).
- Strong command of scraping and data-engineering tooling: headless browsers, HTML and structured-data parsing, scheduling and orchestration, queues, rate-limiting, retries, and monitoring.
- Strong skills in Next.js, React, Supabase/Postgres, Tailwind CSS, and Vercel.
- Responsible AI/LLM integration with cost control and fallbacks.
- UI/UX sensibility with a portfolio of clean, production-grade interfaces.
- Experience preparing iOS and Android apps for App Store and Play Store submission (strongly preferred).
- Fluent English, written and spoken, is mandatory. Clear, proactive communication is essential.
- A disciplined approach to testing and writing sustainable, maintainable, documented code.
ENGAGEMENT TERMS
- One-month, fixed-price contract.
- Paid in milestones: kickoff (after a paid week-one trial task), mid-point on demonstrated progress, and final on acceptance and merge of the completed work.
- A paid trial task in week one, focused on the scraper (for example, adding a new non-trivial source with correct extraction, deduplication, and error handling), is used to confirm fit before fuller commitment. The trial task is paid even if the engagement does not proceed.
- The Consultant must sign a Non-Disclosure Agreement, accept work-for-hire intellectual property terms (all work product owned by ForaHub), work on a separate Git branch, and submit work via pull request for review. No direct pushes or merges to the main branch.
- Access is granted progressively and on a least-privilege basis; production secrets are not shared. All access is revoked at the end of the engagement.
- No destructive actions (deleting data, dropping tables, rotating or exposing secrets, altering production environment variables) without written approval.
- Strong performance may lead to an ongoing or long-term arrangement, agreed separately.
Related categories:
PHP
JavaScript
Python
Web Scraping
PostgreSQL
Next.js
Tailwind CSS
AI Development