Extract Public Legal & Procedural Content from USA Government Website (No Private Data)
Budget: $15 – $25 USD
Extract Public Legal & Procedural Content from USA Government Website (No Private Data)
We are hiring an experienced web scraping professional to extract all publicly available legal, procedural, and educational content from a USA based government website. This project is for internal documentation purposes. NO PRIVATE, personal, or login-protected data is to be accessed.
MAXIMUM PAY: $150 USD
Scope of Work:
Scrape all public-facing HTML content (\~100–120 pages)
Download and extract text from \~30–40 PDF files
Use OCR on any scanned PDFs
Preserve formatting: headings, bullet points, tables
Organize files by category (e.g., FAQs, Forms, Documents, Legal, Guidelines)
Deliver text as .txt or .md files
Provide a metadata log: URL, title, file type, extraction method, word count
Exclusions:
No personal data
No CAPTCHA or login bypass
No scraping of database search tools
No media (videos/images)
Deliverables:
Clean folder of organized text files
Original PDFs
Metadata CSV
Extraction summary report
Project Due:
7 days max
Daily progress updates required.
Screening Question (Required):
SCREENING QUESTION (REQUIRED)
Answer all 3:
1. Tools for HTML and PDF extraction?
2. Have you done legal/government scraping before?
3. How would you handle scanned PDFs with no text layer?
We are hiring an experienced web scraping professional to extract all publicly available legal, procedural, and educational content from a USA based government website. This project is for internal documentation purposes. NO PRIVATE, personal, or login-protected data is to be accessed.
MAXIMUM PAY: $150 USD
Scope of Work:
Scrape all public-facing HTML content (\~100–120 pages)
Download and extract text from \~30–40 PDF files
Use OCR on any scanned PDFs
Preserve formatting: headings, bullet points, tables
Organize files by category (e.g., FAQs, Forms, Documents, Legal, Guidelines)
Deliver text as .txt or .md files
Provide a metadata log: URL, title, file type, extraction method, word count
Exclusions:
No personal data
No CAPTCHA or login bypass
No scraping of database search tools
No media (videos/images)
Deliverables:
Clean folder of organized text files
Original PDFs
Metadata CSV
Extraction summary report
Project Due:
7 days max
Daily progress updates required.
Screening Question (Required):
SCREENING QUESTION (REQUIRED)
Answer all 3:
1. Tools for HTML and PDF extraction?
2. Have you done legal/government scraping before?
3. How would you handle scanned PDFs with no text layer?