Global Argentine Tango Lead Generation & Deduplication
Budget: $250 – $750 USD
I am seeking a high-level Data Engineer to build a comprehensive, deduplicated global database of Argentine Tango events (Festivals, Marathons, and recurring local Milongas). This project is not a simple "copy-paste" task; it requires handling complex JavaScript-heavy sites with infinite scroll and dynamic widgets.
Scope of Work
Data Extraction: Target ~3,000 Festivals/Marathons and ~50,000+ local Milongas worldwide.
Source Management: Scraping from 6-7 global directories (React/JS based) and 3500+ city-specific local hubs.
Technical Categorization: Every entry must be tagged (e.g., Festival, Marathon, Encuentro, or Recurring Milonga).
Data Enrichment: Use Apollo, Clay, or Hunter.io to fill missing organizer emails and WhatsApp numbers.
Advanced Deduplication: Implement a logic where Name + City + Venue + Weekday forms a unique key to merge records from different sources.
Technical Requirements
Stack: Python (Playwright, Selenium, or Scrapy) is preferred to handle "lazy loading" and shadow DOMs.
Logic: Must use Fuzzy Matching (e.g., Levenshtein distance) to merge Source A (Phone) with Source B (Email) without creating duplicates.
Data Quality: Filter for 2025–2026 active dates. No legacy/stale data from 2024 or earlier.
Submission Requirements (Mandatory)
To be considered, your proposal must include:
Technical Stack: Which libraries and "Fuzzy Matching" methods will you use?
The Buenos Aires Test: Provide a 5-row sample (CSV) of active Milongas in Buenos Aires including: Event Name, Venue, Weekday, and a verified WhatsApp/Phone.
Deduplication Plan: Briefly explain how you will handle an event listed on two different sites with slightly different names.
Payment Milestones
Milestone 1: 10% – Successful 100-row pilot (Cleaned & Deduplicated).
Milestone 2: 40% – Full extraction of Global Festivals & Marathons.
Milestone 3: 50% – Final delivery of 30k+ local Milongas with enriched contact data.
Scope of Work
Data Extraction: Target ~3,000 Festivals/Marathons and ~50,000+ local Milongas worldwide.
Source Management: Scraping from 6-7 global directories (React/JS based) and 3500+ city-specific local hubs.
Technical Categorization: Every entry must be tagged (e.g., Festival, Marathon, Encuentro, or Recurring Milonga).
Data Enrichment: Use Apollo, Clay, or Hunter.io to fill missing organizer emails and WhatsApp numbers.
Advanced Deduplication: Implement a logic where Name + City + Venue + Weekday forms a unique key to merge records from different sources.
Technical Requirements
Stack: Python (Playwright, Selenium, or Scrapy) is preferred to handle "lazy loading" and shadow DOMs.
Logic: Must use Fuzzy Matching (e.g., Levenshtein distance) to merge Source A (Phone) with Source B (Email) without creating duplicates.
Data Quality: Filter for 2025–2026 active dates. No legacy/stale data from 2024 or earlier.
Submission Requirements (Mandatory)
To be considered, your proposal must include:
Technical Stack: Which libraries and "Fuzzy Matching" methods will you use?
The Buenos Aires Test: Provide a 5-row sample (CSV) of active Milongas in Buenos Aires including: Event Name, Venue, Weekday, and a verified WhatsApp/Phone.
Deduplication Plan: Briefly explain how you will handle an event listed on two different sites with slightly different names.
Payment Milestones
Milestone 1: 10% – Successful 100-row pilot (Cleaned & Deduplicated).
Milestone 2: 40% – Full extraction of Global Festivals & Marathons.
Milestone 3: 50% – Final delivery of 30k+ local Milongas with enriched contact data.
Related categories:
JavaScript
Python
Web Scraping
Data Mining
Scrapy
Data Extraction
Selenium
API Integration