PII Document Sourcing Specialist
Budget: ₹100 – ₹400 INR
I am assembling a corpus of real-world documents that contain publicly available PII across multiple domains and need a consultant who can combine smart desk-research with targeted web-scraping to locate, capture, and catalogue them for later analysis.
Scope of work
• Source material: real-world data sets, reports, filings, or any document that actually exposes PII (names, addresses, SSNs, medical record numbers, account details, etc.).
• Focus domains: Healthcare, Finance and Education.
• Geographic emphasis: North-American publications at this stage (government portals, open-data sites, public court documents, regulatory disclosures, etc.); other regions may follow once this tranche is complete.
• Methods: a mix of automated scraping (Python, BeautifulSoup/Scrapy/Selenium or similar) and classic desk research to reach sources that resist automation.
• Output: an organised folder structure plus a spreadsheet/JSON catalog listing document title, source URL, date accessed, domain tag, and a short note of the specific PII fields present.
Acceptance criteria
1. Minimum 250 unique documents, balanced across the three domains and all the documents ya must be single page with maximum 300 words
2. Each entry must include working source links and clear evidence of at least one PII field.
3. No paywalled or illegally obtained content—everything must be freely reachable on the open web.
4. Scripts (if used) are handed over, well-commented, and runnable on a vanilla Python environment.
If you have experience harvesting open data, navigating government portals, and keeping an eye on compliance while still finding those hard-to-spot files, I’d like to hear how you’d tackle this and how quickly you could deliver the first batch.
Scope of work
• Source material: real-world data sets, reports, filings, or any document that actually exposes PII (names, addresses, SSNs, medical record numbers, account details, etc.).
• Focus domains: Healthcare, Finance and Education.
• Geographic emphasis: North-American publications at this stage (government portals, open-data sites, public court documents, regulatory disclosures, etc.); other regions may follow once this tranche is complete.
• Methods: a mix of automated scraping (Python, BeautifulSoup/Scrapy/Selenium or similar) and classic desk research to reach sources that resist automation.
• Output: an organised folder structure plus a spreadsheet/JSON catalog listing document title, source URL, date accessed, domain tag, and a short note of the specific PII fields present.
Acceptance criteria
1. Minimum 250 unique documents, balanced across the three domains and all the documents ya must be single page with maximum 300 words
2. Each entry must include working source links and clear evidence of at least one PII field.
3. No paywalled or illegally obtained content—everything must be freely reachable on the open web.
4. Scripts (if used) are handed over, well-commented, and runnable on a vanilla Python environment.
If you have experience harvesting open data, navigating government portals, and keeping an eye on compliance while still finding those hard-to-spot files, I’d like to hear how you’d tackle this and how quickly you could deliver the first batch.
Related categories:
Python
Research
Web Scraping
Data Mining
Data Extraction
Data Collection
Data Management
Document Checking