Model PDF Dataset Collection
Budget: $10 – $30 USD
I need every PDF listed on the two pages below downloaded and sorted into a clean, intuitive folder tree:
https://tosotdirect.com/pages/product-document
https://www.greecomfort.com/system-documentation/
Folder structure must be exactly:
brand / model / type / file.pdf
• brand = site‐listed brand name
• model = product model shown in the listing
• type = the document category (e.g., user-manual, technical-specs) as labelled on the site
• file.pdf = keep the original file name—no renaming, no re-ordering.
A simple ZIP of the finished dataset is fine.
I’ll review by unpacking the archive and confirming:
1. Every PDF on the two pages is present.
2. The directory path follows the brand/model/type pattern.
3. Filenames are untouched.
No additional sites are needed for this round, though I may extend the project later. A quick script, scraper, or manual download—use whatever method works best, as long as the final folders match the spec.
Deliverables:
• Zipped dataset containing all PDFs in the required hierarchy.
• Short note (TXT or Markdown) describing the steps or script used, so the process can be repeated later if needed.
That’s it—straightforward collection and organization.
https://tosotdirect.com/pages/product-document
https://www.greecomfort.com/system-documentation/
Folder structure must be exactly:
brand / model / type / file.pdf
• brand = site‐listed brand name
• model = product model shown in the listing
• type = the document category (e.g., user-manual, technical-specs) as labelled on the site
• file.pdf = keep the original file name—no renaming, no re-ordering.
A simple ZIP of the finished dataset is fine.
I’ll review by unpacking the archive and confirming:
1. Every PDF on the two pages is present.
2. The directory path follows the brand/model/type pattern.
3. Filenames are untouched.
No additional sites are needed for this round, though I may extend the project later. A quick script, scraper, or manual download—use whatever method works best, as long as the final folders match the spec.
Deliverables:
• Zipped dataset containing all PDFs in the required hierarchy.
• Short note (TXT or Markdown) describing the steps or script used, so the process can be repeated later if needed.
That’s it—straightforward collection and organization.
Related categories:
PHP
JavaScript
Python
Web Scraping
Software Architecture
Scrapy
BeautifulSoup
Data Collection