One-Month Website Text Extraction

Job ID: 40433160

Budget: €12 – €18 EUR

I have a collection of websites whose written content needs to be captured regularly over the coming month. The task is straightforward: pull all required text, keep its structure intact (titles, sub-headings, body copy, meta information when available), and funnel the results into a clean Excel workbook I can sort and filter with ease.

I’ll confirm the exact URLs and the specific page sections to target before we start. From there, I’d like an initial sample run so we can agree on field order, encoding, and any pre-processing rules such as stripping HTML tags or preserving line breaks. Once that sample is approved, the same routine should run on a schedule we agree upon—daily is preferred—until the full month is complete.

You may use whichever tools you’re comfortable with (Python, Scrapy, BeautifulSoup, Selenium, or similar) as long as the final output is a consistently formatted .xlsx file and the scraping respects each site’s robots.txt and rate limits.

Deliverables:
• Approved sample spreadsheet for one site
• Daily or weekly Excel files for all targeted sites
• A brief note of any anomalies or access issues encountered during the month

I’m ready to start immediately and will be responsive to settle details quickly.