Python Scraper Engineer – Cloudflare Bypass & Redis Session Mgmt

Job ID: 40571572

Budget: $2 – $8 USD

We're looking for an experienced Python developer to stabilize and enhance a production web scraper that's experiencing Cloudflare blocks, broken session handling, and WebSocket instability. The goal is to implement a robust, self-healing scraping pipeline with proper Redis integration.

Requirements:
- Strong Python experience with web scraping frameworks (Playwright, curl-cffi, Scrapy, or similar)
- Hands-on experience bypassing Cloudflare challenges using TLS fingerprint matching and stealth techniques
- Redis experience including session storage, TTL management, and pub-sub patterns
- Experience with job queuing libraries such as BullMQ or RQ
- WebSocket client implementation including reconnection logic, heartbeat management, and binary/JSON frame parsing
- Ability to integrate into an existing codebase without unnecessary rewrites

Deliverables:
- Cloudflare bypass layer using curl-cffi with browser-matched TLS fingerprints and playwright-stealth for JS challenge handling
- Redis-backed session store with cookie expiry tracking and automatic re-authentication on stale sessions
- Intelligent session rotation logic that retires sessions before flagging, not after
- Persistent WebSocket client with exponential backoff reconnection and proper heartbeat management
- BullMQ or RQ job queue integration with retry logic and scraped data caching
- Clean integration into the existing scraper codebase with minimal disruption

Work Arrangement:
This is an hourly engagement. We expect ongoing collaboration as we work through each component of the pipeline. Availability to start quickly is a plus.

About the project:
This is an active production scraper that needs surgical improvements to reduce block rates, eliminate session crashes, and stabilize WebSocket connections — reliability and uptime are the top priorities.