Wanted: Developer to Build Web Scraping & Data-Update Monitoring System

Job ID: 39650952

Budget: $15 – $25 AUD

I’m seeking an experienced developer to build an end‑to‑end automated data‑scraping system focused on nutritional information from restaurant websites and menu PDFs. This system must be robust, scalable, and easy for me to maintain once it’s operational.

Core features and tasks:

• Data collection: Scrape nutritional tables and menu data from a list of restaurant websites and downloadable PDF files (which may be formatted differently across sites). If pages use JavaScript to load content or contain CAPTCHAs, you’ll need to handle those appropriately.

• Data extraction and normalisation: Extract relevant fields (e.g. product names, serving sizes, calories, fats, sugars, allergens) and normalise them into a consistent schema across all sources.

• Storage and formatting: Store the scraped data in a structured format, preferably in spreadsheets (Google Sheets or Excel) but open to a simple database if it ensures better integrity and scalability. Each run should save a new snapshot to facilitate change tracking.

• Change monitoring: Compare new runs with previous data to detect additions, deletions, or modifications. Generate a clear log of differences.

• Notifications: Notify me of any changes via email and possibly Slack or a lightweight dashboard (e.g. a simple web interface) showing the latest updates and historical changes.

• Scheduling and automation: Schedule the scraping and update‑checking to run automatically on a daily or weekly basis (Sydney/Australia timezone). The system should be resilient to layout changes on target sites, and it should be easy to update the list of sites/PDFs to monitor.

• Documentation & handover: Deliver well‑commented code, clear documentation, and setup instructions so that I can run and maintain the system without needing to rehire.

Ideal Skills and Experience

• Web scraping expertise: Experience using tools/libraries such as Python with BeautifulSoup, Scrapy, Selenium, or Playwright to handle static and dynamic websites, including basic anti‑bot mechanisms.

• Data extraction & normalisation: Proven ability to parse tables and unstructured text from both HTML and PDF documents, then organise them into a unified structure.

• Spreadsheet/database knowledge: Comfortable storing data in spreadsheets (Google Sheets API or CSV/Excel) and/or lightweight databases while ensuring traceability and versioning.

• Change detection & alerts: Familiarity with methods for comparing datasets across runs and triggering notifications (email, Slack, webhook) when new or changed entries appear.

• Automation & scheduling: Experience setting up cron jobs or scheduled tasks on a server (or cloud functions) to run reliably in my timezone and recover gracefully from errors or site changes.

• Clean coding & documentation: Ability to write maintainable code with clear documentation, unit tests, and deployment instructions.

Please include in your bid:

• A short description of similar scraping/monitoring projects you’ve done, particularly involving PDFs or structured nutritional data.

• Your proposed approach and timeline to complete this project, including any questions you need answered up front.

• Whether you’re comfortable working in Australia/Sydney timezone for scheduling the tasks and possible follow‑up.

Looking forward to reviewing your proposals!