Apify Specialist for Targeted Data Extraction
Budget: $30 – $250 CAD
I am looking for an experienced Apify specialist to execute a highly targeted, small-scale data extraction project. The goal is to collect 200–300 high-signal records from specific professional discussion forums and community platforms.
This is not a complex web development project. I have already created a an execution plan. Your job is to configure the tools in my workspace, execute the runs cleanly, and normalize the exported data.
Scope of Work:
Task Setup: Configure specific Apify actors (primarily Web Scraper and Reddit scrapers) directly within my Apify workspace so I retain the assets.
Execution: Run pre-defined search queries (which will be provided upon hire) across multiple platforms.
Strict Filtering: Apply post-run filters to the dataset. For example, on certain platforms, you must filter the dataset to only include records where the author possesses specific verified credential flairs.
Data Normalization: Clean the raw export before final delivery. This includes deduplicating records and mapping synonym terms to a canonical name (e.g., mapping various abbreviations to a single standard name).
Delivery: Provide the final cleaned dataset in both JSON and CSV formats.
Target Sources:
Reddit (specific subreddits, filtering for verified credentials)
Doximity / OpMed
Medscape
HealthUnlocked
Pharmacy Times
Requirements & Milestones:
To ensure quality and alignment, this project will be strictly managed via milestones. Do not bid if you cannot agree to the following workflow:
Milestone 1: Delivery and verification of Tier 1 sources (~100 records).
Milestone 2: Delivery and verification of Tier 2 sources (~100-150 records) and final data normalization.
To Apply:
Please start your proposal with the word "CANONICAL" so I know you read this entire posting. In your proposal, briefly confirm your experience with Apify and state how you handle data normalization (e.g., Excel, Python, Pandas) for the final CSV delivery.
This is not a complex web development project. I have already created a an execution plan. Your job is to configure the tools in my workspace, execute the runs cleanly, and normalize the exported data.
Scope of Work:
Task Setup: Configure specific Apify actors (primarily Web Scraper and Reddit scrapers) directly within my Apify workspace so I retain the assets.
Execution: Run pre-defined search queries (which will be provided upon hire) across multiple platforms.
Strict Filtering: Apply post-run filters to the dataset. For example, on certain platforms, you must filter the dataset to only include records where the author possesses specific verified credential flairs.
Data Normalization: Clean the raw export before final delivery. This includes deduplicating records and mapping synonym terms to a canonical name (e.g., mapping various abbreviations to a single standard name).
Delivery: Provide the final cleaned dataset in both JSON and CSV formats.
Target Sources:
Reddit (specific subreddits, filtering for verified credentials)
Doximity / OpMed
Medscape
HealthUnlocked
Pharmacy Times
Requirements & Milestones:
To ensure quality and alignment, this project will be strictly managed via milestones. Do not bid if you cannot agree to the following workflow:
Milestone 1: Delivery and verification of Tier 1 sources (~100 records).
Milestone 2: Delivery and verification of Tier 2 sources (~100-150 records) and final data normalization.
To Apply:
Please start your proposal with the word "CANONICAL" so I know you read this entire posting. In your proposal, briefly confirm your experience with Apify and state how you handle data normalization (e.g., Excel, Python, Pandas) for the final CSV delivery.
Related categories:
C Programming
Data Processing
Web Scraping
Data Mining
JSON
Data Extraction
API
Data Analysis
Data Collection
Data Management