Daily Web Scraping for Market Research
Budget: $250 – $750 USD
I’m looking for a data engineer who can take full ownership of a daily web-scraping workflow aimed at ongoing market research. The job centers on extracting selected data points from public web pages, transforming them into a clean, structured format, and making them available for analysis every 24 hours.
Here’s what I need you to handle from end to end:
• Source acquisition – fetch HTML from the URLs I provide, even when content is hidden behind JavaScript (a headless browser such as Playwright or Selenium is fine).
• Parsing & cleansing – pull the specific fields I’ll list (product name, price, SKU, availability, and a time-stamp), remove duplicates, and standardize values.
• Storage & delivery – load the daily output into my PostgreSQL instance; if you prefer Parquet or plain CSV that’s acceptable as long as it’s automated.
• Orchestration – schedule the run with Cron, Airflow, or an equivalent tool, include retry logic, logging, and simple alerting on failure.
• Compliance – respect robots.txt where applicable and rotate proxies or user agents to avoid blocks.
Acceptance criteria
1. A Git-based repository containing well-commented code and a README with setup instructions.
2. One successful live run on my infrastructure demonstrating the full daily cycle.
3. Data accuracy ≥ 98 % on the agreed fields, verified against a manual sample.
If you can outline your preferred stack, past experience scraping web pages for market-research use cases, and any ideas for improving reliability, we can move forward quickly.
Here’s what I need you to handle from end to end:
• Source acquisition – fetch HTML from the URLs I provide, even when content is hidden behind JavaScript (a headless browser such as Playwright or Selenium is fine).
• Parsing & cleansing – pull the specific fields I’ll list (product name, price, SKU, availability, and a time-stamp), remove duplicates, and standardize values.
• Storage & delivery – load the daily output into my PostgreSQL instance; if you prefer Parquet or plain CSV that’s acceptable as long as it’s automated.
• Orchestration – schedule the run with Cron, Airflow, or an equivalent tool, include retry logic, logging, and simple alerting on failure.
• Compliance – respect robots.txt where applicable and rotate proxies or user agents to avoid blocks.
Acceptance criteria
1. A Git-based repository containing well-commented code and a README with setup instructions.
2. One successful live run on my infrastructure demonstrating the full daily cycle.
3. Data accuracy ≥ 98 % on the agreed fields, verified against a manual sample.
If you can outline your preferred stack, past experience scraping web pages for market-research use cases, and any ideas for improving reliability, we can move forward quickly.
Related categories:
Python
Data Processing
Web Scraping
Data Mining
PostgreSQL
Elasticsearch
Data Analysis
Selenium