Directory Web Scraping
Budget: $100 – $110 USD
We are looking for an experienced web scraping specialist to help us extract publicly available contact data from an educational institution directory , located at the following URL:
The site lists individual profile pages for each institution (academy or training center), accessible through unique numeric IDs. We want to collect data from all valid and accessible profiles.
Data to Extract:
For each institution profile, we require the following information:
Name of the institution
Email address
Website URL
Instagram or other social media links (if available)
Phone number
Location (city/province, if available)
Direct link to the profile page
All data should be delivered in a structured CSV or Excel file.
Requirements:
Proven experience in web scraping
Ability to iterate through numeric IDs or handle pagination
Implement error handling and retry logic for missing or invalid pages
Skill in parsing text to extract emails, phone numbers, and URLs using regular expressions or similar techniques
Respectful of server limits (apply appropriate delays between requests)
Provide clean, well-commented code and basic documentation
Deliverables:
A CSV or Excel file with the complete and organized dataset
Bonus Points:
Ability to set up automation or periodic scraping
Ability to detect and extract social media links within unstructured text
Optionally clean or validate extracted data before delivery
The site lists individual profile pages for each institution (academy or training center), accessible through unique numeric IDs. We want to collect data from all valid and accessible profiles.
Data to Extract:
For each institution profile, we require the following information:
Name of the institution
Email address
Website URL
Instagram or other social media links (if available)
Phone number
Location (city/province, if available)
Direct link to the profile page
All data should be delivered in a structured CSV or Excel file.
Requirements:
Proven experience in web scraping
Ability to iterate through numeric IDs or handle pagination
Implement error handling and retry logic for missing or invalid pages
Skill in parsing text to extract emails, phone numbers, and URLs using regular expressions or similar techniques
Respectful of server limits (apply appropriate delays between requests)
Provide clean, well-commented code and basic documentation
Deliverables:
A CSV or Excel file with the complete and organized dataset
Bonus Points:
Ability to set up automation or periodic scraping
Ability to detect and extract social media links within unstructured text
Optionally clean or validate extracted data before delivery