? Scraping Task – Extracting Emails from Institutions (France)

Job ID: 39181493

Budget: €8 – €30 EUR

Objective: Extract only email addresses from two websites in France.

Task:
✅ Scrape emails from two websites
✅ Remove duplicates after extraction
✅ Deliver a clean Excel file with unique emails

Details of the Websites to Scrape
Site 1
https://trajectoire.sante-ra.fr/Trajectoire/pages/AccesLibre/Annuaires/EtablissementMS.aspx

Instructions:
1️⃣ Enter the postal code → 03100 – MONTLUÇON
2️⃣ Select the distance → Illimitée
3️⃣ Click on "Rechercher"
4️⃣ A list of institutions with images will appear → Click on each image to open the profile page.
5️⃣ Scroll to the bottom of the profile to find one or more email addresses.
6️⃣ Extract only the emails.

✅ Total records to process: 9,355

Site 2
https://www.pour-les-personnes-agees.gouv.fr/acces-annuaires?annuaire=GENERAL#je-recherche-par-annuaire

Instructions:
1️⃣ Select all 6 institution categories:

Accueil de jour
Établissement d'hébergement pour personnes âgées dépendantes (EHPAD)
Établissement d'hébergement pour personnes âgées (EHPA)
Service autonomie à domicile (aide)
Service autonomie à domicile (aide et soins)
Unité de soins de longue durée (USLD)
2️⃣ Search by department:

France is divided into 100 departments (01 – Ain to 95 – Val d’Oise)
Plus 971, 972, 973, 974, 976 (French overseas territories).
Enter "01", select Ain (01), then click on "Rechercher"
3️⃣ A list of institutions will appear → Click on "Contact"
4️⃣ The system will open the email address directly in an email client.
5️⃣ Extract only the emails.

✅ Estimated total records: around 30,000 to 50,000

Total to process:
~40,000 to 60,000 records → Expected: 50,000 to 60,000 emails

Delivery Format:
✔ Separate files for each website
✔ For Site 2: emails organized by category
✔ A final file containing all deduplicated email addresses

Deadline: March 16, 2024

Thank you!
Related categories: Data Entry Excel Web Scraping Web Search Data Mining