Hardcore Scraper Needed — 350 Mixed Domains (Playwright/Proxies/Anti-Bot)
Budget: $250 – $750 USD
I need a hardcore scraper to extract property listings from ~350 unique URLs (different domains). This is not a simple BeautifulSoup job — many sites are behind Cloudflare, DataDome, hCaptcha, and heavy JavaScript.
I already have a Python skeleton script (scrape_master.py). You may:
Use and improve my script (integrate ScrapFly, Playwright, proxies, captcha-solving, sticky sessions), OR
Build your own solution — but you must run it yourself and deliver the results.
Do not ask me to run the script. You run it, scrape the data, and deliver clean CSVs.
Technical requirements:
First try requests; fallback to Playwright/Puppeteer with stealth + proxies if HTML is blocked or empty.
Handle Cloudflare, DataDome, hCaptcha challenges (ScrapFly, DeathByCaptcha, proxies, sticky sessions).
Maintain cookies/session per domain when required.
Record failures with reason in manual_review.csv.
Provide logs in probe_results.csv (URL, status, HTML length, notes).
Data fields to extract (from text + structured data):
Title
Content (description)
Images
Address
Latitude / Longitude
Bedrooms, Bathrooms
Size (convert SqFt → m²)
Price + Currency
Listing URL
Agency
ISO country code
Primary image
Features (interior + exterior details from description text: e.g., pool, terrace, garden, parking, furnished/unfurnished, etc.)
Deliverables:
properties_import.csv (with all fields above)
agents_import.csv (agent name, phone, WhatsApp, social links, role, description)
enriched_agencies.csv + profile_import.csv (if applicable)
probe_results.csv + manual_review.csv
Updated script(s) + short README (how you ran it, services used, estimated cost of ScrapFly/DeathByCaptcha)
Acceptance criteria:
≥85% of URLs scraped successfully OR logged in manual_review.csv with clear reason.
Data normalized (SqFt → m²).
Interior/exterior features extracted from free text where possible.
Script runs on your side; I get CSV outputs + code.
Test phase (paid):
Run on 20 mixed URLs (I provide).
Deliver all CSVs + short report.
If test succeeds, full project (350 URLs).
Budget & timeline:
Fixed price. Quote for the 20-URL test and for the full 350 URLs.
Delivery for test: 3 days.
I already have a Python skeleton script (scrape_master.py). You may:
Use and improve my script (integrate ScrapFly, Playwright, proxies, captcha-solving, sticky sessions), OR
Build your own solution — but you must run it yourself and deliver the results.
Do not ask me to run the script. You run it, scrape the data, and deliver clean CSVs.
Technical requirements:
First try requests; fallback to Playwright/Puppeteer with stealth + proxies if HTML is blocked or empty.
Handle Cloudflare, DataDome, hCaptcha challenges (ScrapFly, DeathByCaptcha, proxies, sticky sessions).
Maintain cookies/session per domain when required.
Record failures with reason in manual_review.csv.
Provide logs in probe_results.csv (URL, status, HTML length, notes).
Data fields to extract (from text + structured data):
Title
Content (description)
Images
Address
Latitude / Longitude
Bedrooms, Bathrooms
Size (convert SqFt → m²)
Price + Currency
Listing URL
Agency
ISO country code
Primary image
Features (interior + exterior details from description text: e.g., pool, terrace, garden, parking, furnished/unfurnished, etc.)
Deliverables:
properties_import.csv (with all fields above)
agents_import.csv (agent name, phone, WhatsApp, social links, role, description)
enriched_agencies.csv + profile_import.csv (if applicable)
probe_results.csv + manual_review.csv
Updated script(s) + short README (how you ran it, services used, estimated cost of ScrapFly/DeathByCaptcha)
Acceptance criteria:
≥85% of URLs scraped successfully OR logged in manual_review.csv with clear reason.
Data normalized (SqFt → m²).
Interior/exterior features extracted from free text where possible.
Script runs on your side; I get CSV outputs + code.
Test phase (paid):
Run on 20 mixed URLs (I provide).
Deliver all CSVs + short report.
If test succeeds, full project (350 URLs).
Budget & timeline:
Fixed price. Quote for the 20-URL test and for the full 350 URLs.
Delivery for test: 3 days.
Related categories:
PHP
JavaScript
Python
Web Scraping
Scrapy
Data Extraction
API
BeautifulSoup
Automation
API Development