Scrape 54M PDFs from ReJust

Job ID: 40481732

Budget: $5,000 – $10,000 USD

I need a complete, verifiable archive of every PDF that can be generated from the rejust.ro website—roughly fifty-four million documents in total. For each successful download I will pay 0.1 ¢, which brings the full payout to about $5,400 once the entire set is delivered.

Organisation & naming
• Save the files under a clear directory structure of your choice, as long as it reflects a logical category scheme (for example, first-letter buckets or date ranges) and keeps filenames human-readable.
• Each individual PDF must be named with the page title pulled directly from its source page on rejust.ro. That title is the definitive keyword; no incremental numbering or hashes in the filename.

Acceptance criteria
• All accessible pages on rejust.ro have been processed and every resulting PDF is present.
• Filenames match the exact page titles without truncation (UTF-8-safe).
• Folder hierarchy is consistent and reproducible.
• A simple log or database lists URL → saved path for spot-checking.
• Anything that fails on first pass must be re-queued automatically so the final count equals the site’s total.

Technical notes
I expect you will automate this with something robust—Python + Scrapy, Playwright, or similar—and host interim data in cloud storage (AWS S3, GCS, etc.) so progress can be monitored and partial batches handed over if needed. Please factor in throttling, retries and resume capability; the crawl should be polite yet efficient.

When you bid, outline:
1. Your harvesting approach and tooling.
2. Estimated runtime and infrastructure you’ll use.
3. How you plan to hand off such a large data set (drive shipment, cloud bucket, etc.).

A quick proof-of-concept—say, the first 500 PDFs—will secure the contract so we can scale up to the entire corpus immediately.

If auto cannot be done I am okay with manual downloading.