Website Scrape to Word Docs
Budget: $30 – $250 USD
I need the full content of my website—both the text and every image—captured and rebuilt inside Microsoft Word. Please treat each existing section or category on the site as its own document so the final set feels organised rather than stitched together.
Formatting is important:
• Headings and sub-headings must reflect the site hierarchy
• Bullet points or numbered lists should match what’s online
• Text and images have to sit inline so the flow reads naturally
When you pull the images, keep their quality and any captions or alt-text intact. If Word strips metadata, include a separate, well-named image folder alongside the documents.
Deliverables
1. One Word file per section/category, fully formatted as above
2. Supporting image folder (only if Word removes the originals’ data)
3. A short log noting any page you couldn’t scrape and the reason
The site is publicly accessible, so no login hurdles. Let me know which tools (Python, BeautifulSoup, Scrapy, browser automation, etc.) you prefer and how quickly you can turn this around.
Formatting is important:
• Headings and sub-headings must reflect the site hierarchy
• Bullet points or numbered lists should match what’s online
• Text and images have to sit inline so the flow reads naturally
When you pull the images, keep their quality and any captions or alt-text intact. If Word strips metadata, include a separate, well-named image folder alongside the documents.
Deliverables
1. One Word file per section/category, fully formatted as above
2. Supporting image folder (only if Word removes the originals’ data)
3. A short log noting any page you couldn’t scrape and the reason
The site is publicly accessible, so no login hurdles. Let me know which tools (Python, BeautifulSoup, Scrapy, browser automation, etc.) you prefer and how quickly you can turn this around.
Related categories:
Python
Data Processing
Web Scraping
Scrapy
Data Extraction
BeautifulSoup
Automation
Microsoft Word