Directory Text & Image Scraping

Job ID: 39843816

Budget: $10 – $30 USD

I need a clean, reliable script that will crawl a web-based directory, pull down every piece of text and each associated image, then organise everything so I can analyse it offline. The directory entries hold the bulk of the text, and for every listing I need the contact information and the core business details captured alongside any images displayed on the page.

A typical record in the final file should show fields such as business name, address, phone, email, website and a brief description, with the corresponding image either stored locally (and referenced by filename) or recorded as its direct URL. Accuracy is vital—no duplicates, no half-filled rows.

You are free to choose your preferred stack; Python with BeautifulSoup, Scrapy or Selenium is perfectly fine, as is any other tool that gets the job done quickly and safely without breaching the site’s usage terms. Make sure the script is well-commented so I can run it again later if the directory updates.

Deliverables:
• Executable or runnable script with clear setup instructions
• Structured CSV or Excel file containing all directory data
• Folder or archive of the scraped images (or a list of live URLs)
• Brief read-me outlining how to rerun or tweak the crawl

Once everything is tested and spot-checked for completeness, the project is complete.