Weekly Daft.ie MyHome Scraper
Budget: $30 – $250 USD
I need a person to create and manage an automated script that once a week sweeps both daft.ie and myhome.ie, grabs every new or updated property, and stores the full asking price, complete text description, and all available photos so I can revisit the brochure months after the original ad has vanished.
I will pay a set up and small monthly management fee of the process
Data format isn’t fixed—CSV, Excel, JSON or anything similarly straightforward all work as long as I can open it later. What matters is that the archive is tidy and searchable.
Key deliverables
• A dependable scraper (Python, Scrapy/Selenium/BeautifulSoup—your call) scheduled to run weekly without manual input.
• De-duping logic based on listing ID or URL so the same property is never saved twice.
• For each listing, a folder containing the pictures plus a metadata file holding price and description.
• Initial back-fill of current live listings so the archive starts complete.
• Basic logging so I can quickly confirm when the job last ran and how many ads were captured.
Acceptance: after the first scheduled run I should be able to open a saved folder, see the images offline, and read the exact price and wording that originally appeared online.
Please outline any anti-scraping precautions you’ll take (rate limiting, rotating headers, etc.), the stack you prefer, and how long you need to deliver the first working run.
I will pay a set up and small monthly management fee of the process
Data format isn’t fixed—CSV, Excel, JSON or anything similarly straightforward all work as long as I can open it later. What matters is that the archive is tidy and searchable.
Key deliverables
• A dependable scraper (Python, Scrapy/Selenium/BeautifulSoup—your call) scheduled to run weekly without manual input.
• De-duping logic based on listing ID or URL so the same property is never saved twice.
• For each listing, a folder containing the pictures plus a metadata file holding price and description.
• Initial back-fill of current live listings so the archive starts complete.
• Basic logging so I can quickly confirm when the job last ran and how many ads were captured.
Acceptance: after the first scheduled run I should be able to open a saved folder, see the images offline, and read the exact price and wording that originally appeared online.
Please outline any anti-scraping precautions you’ll take (rate limiting, rotating headers, etc.), the stack you prefer, and how long you need to deliver the first working run.
Related categories:
JavaScript
Python
Data Processing
Web Scraping
Software Architecture
Scrapy
BeautifulSoup
Selenium
Automation
Data Management