Extract Public Legal & Procedural Content from USA Government Website (No Private Data)

Job ID: 39504208

Budget: $15 – $25 USD

Extract Public Legal & Procedural Content from USA Government Website (No Private Data)
We are hiring an experienced web scraping professional to extract all publicly available legal, procedural, and educational content from a USA based government website. This project is for internal documentation purposes. NO PRIVATE, personal, or login-protected data is to be accessed.

MAXIMUM PAY: $150 USD

Scope of Work:

Scrape all public-facing HTML content (\~100–120 pages)
Download and extract text from \~30–40 PDF files
Use OCR on any scanned PDFs
Preserve formatting: headings, bullet points, tables
Organize files by category (e.g., FAQs, Forms, Documents, Legal, Guidelines)
Deliver text as .txt or .md files
Provide a metadata log: URL, title, file type, extraction method, word count

Exclusions:

No personal data
No CAPTCHA or login bypass
No scraping of database search tools
No media (videos/images)

Deliverables:

Clean folder of organized text files
Original PDFs
Metadata CSV
Extraction summary report

Project Due:
7 days max

Daily progress updates required.

Screening Question (Required):
SCREENING QUESTION (REQUIRED)
Answer all 3:

1. Tools for HTML and PDF extraction?
2. Have you done legal/government scraping before?
3. How would you handle scanned PDFs with no text layer?