AI-Driven Kids’ Bike Database Updater (Delivery by 31 Aug 2025)

Job ID: 39671367

Budget: €750 – €1,500 EUR

Project Overview
We need an end-to-end solution that monitors kids’ bicycle listings across 50-plus manufacturer sites and 20-plus e-commerce shops, then inserts, updates, or deactivates records in our PostgreSQL database. A rule-based “Phase 1” proof quickly validates the approach on one source; subsequent phases scale to all sites. All code must run on our Windows Server with straightforward logging and scheduling.

You’ll receive on award:
A slide deck outlining goals, workflow and KPIs.
The target DB schema (PostgreSQL, about 20 tables).
CSV lists of 48 OEM sites and 23 retailers to scrape.

Scope & Milestones:
Phase 1 – Pilot: Build a rule-based classifier to tag kids’ bikes, scrape one manufacturer site end-to-end, and write/update records in a staging PostgreSQL database. Acceptance: at least 95 % product-match accuracy on the pilot site and automatic deactivation of discontinued SKUs.

Phase 2 – OEM Scrapers: Develop scrapers for all 48 manufacturer websites, extract core specs (wheel size, frame material, MSRP, etc.), and add robust retry and change-detection logic. Acceptance: a daily cron run finishes in under two hours with an overall error rate below 2 %.

Phase 3 – Web-Shop Scrapers: Implement scrapers for 23 retailer sites (handling pagination and price variants) and map retailer listings back to their corresponding OEM models. Acceptance: at least 90 % coverage of retailer links stored, with retail prices correctly captured.

Phase 4 – Deployment & Handover: Provide CI/CD scripts, a comprehensive README, and conduct a one-hour knowledge-transfer call. Acceptance: all automated tests pass in our server environment and the system runs reliably on schedule.

We’re open to revising the milestone plan if you propose a better approach.

Tech Stack:
Python 3.11 (Scrapy, Playwright or Selenium for dynamic sites).
PostgreSQL, SQLAlchemy / Psycopg.
Pandas (ETL), simple rules or lightweight ML for bike detection.
Docker (optional)

What We’re Looking For:
Proven track record scraping complex, JS-heavy sites at scale.
Clean, well-documented, test-covered code.
Ability to hit a hard deadline of 31 Aug 2025.
Excellent communicator (weekly demo + Slack updates).

To Apply
Briefly outline your approach (libraries, anti-bot strategy, scheduling).
Share relevant portfolio links (GitHub, past scraping/ETL projects).
State estimated budget (fixed or milestone-based) and earliest start date.
List any questions or assumptions you need clarified.

Important Notes
The attached deck, schema, and site lists are for shortlisted bidders only (NDA required).
We need daily progress notes once work begins.
Final deliverable must be easy for another developer to maintain.
Looking forward to your creative solutions—let’s build something great together!