Web Scraper Development for Leaderboards.

Job ID: 40210783

Budget: $10 – $30 USD

Please Read Carefully Before Applying
It does not matter whether you consider yourself a “vibe coder” or a traditional software engineer we accept both here.
What matters is whether you can make this system work reliably at scale.
We operate a production scraper that processes 500+ leaderboard sites per hour.
All sites we scrape are leaderboards, but no two sites are the same.
This is not a basic scraper.
What Makes This Scraper Different
The leaderboards we scrape vary heavily in structure and behavior:
Dynamic buttons, tabs, and switchers
JavaScript-rendered content
Hybrid navigation (UI interaction + background API calls)
Tables, card layouts, podium layouts, or combinations of all three
Masked usernames and inconsistent rank formats
Different ordering of wager / prize data
Many sites require:
Clicking the correct UI elements to reveal data
Detecting when the full leaderboard is loaded
Handling asynchronous data sources
Choosing between multiple conflicting data representations
This requires a hybrid scraping approach, not a static script.
Current State of the System
We already have a working base scraper that was partially vibe-coded.
It performs well on many sites, but it is not yet robust enough at large scale.
Current limitations include:
Inconsistent detection of leaderboard buttons and tabs
Partial leaderboards being captured instead of full datasets
Site-specific edge cases breaking otherwise valid logic
Some extraction methods working well in isolation but failing in combination
Your role is to improve reliability, coverage, and correctness, not to rebuild everything.
Technologies & Tools in Use
These are important to note before applying:
Node.js
Playwright for browser automation
Network interception for API-based extraction
HTML → Markdown conversion for structural parsing
OCR-based fallback extraction for edge cases
PostgreSQL + TimescaleDB for storage
Cron-based execution with hourly runs
The system uses multiple extraction methods in parallel and merges results based on confidence and validation rules.
AI & Vision-Based Testing
We have performed internal safety and validation testing using Claude Vision, particularly for:
Visual leaderboard layout interpretation
Cross-checking OCR-based extraction accuracy
Identifying layout patterns that are difficult to detect via DOM alone
This testing has shown promising results, but it is used as a supporting validation tool, not a replacement for deterministic scraping logic.
Experience or interest in vision-assisted validation is a plus, but not required.
What We’re Looking For
We are looking for a serious and capable person who:
Understands real-world web scraping challenges
Can reason through inconsistent, dynamic site behavior
Implements practical, scalable solutions
Focuses on reliability, performance, and correctness
Whether you “vibe code” or write highly structured code does not matter.
If it works at scale, it’s acceptable.
Scale & Performance Expectations
This scraper must:
Handle hundreds of sites per run, we'll provide a list of sites in a txt.
Execute in a timely and predictable manner

Save and update data efficiently

Support near-live internal updates

Solutions that only work for a handful of sites are not sufficient.

Before Applying

Please do not apply unless you have read and understood this post in full.

We are not looking for:
Trial-and-error scrapers
One-size-fits-all assumptions
People unfamiliar with dynamic or JS-heavy sites
We are looking for someone who can actually solve this problem.
Price is non-negotiable, so please bid if you feel this price is stable.
I will circle back as soon as possible.