Website Text Scraping to Excel
Budget: €8 – €30 EUR
I need every piece of text that appears in the site’s tables and list elements captured and delivered in a neatly organised Excel workbook. Once the crawl is complete, each row or list item should occupy its own row in the spreadsheet, with column headers that clearly label every field you have pulled.
Accuracy is key: the sheet must mirror the live site, without missing or duplicated entries. If a page is paginated, please make sure all pages are traversed. Dynamic content that loads as you scroll should be handled as well.
A lightweight, well-commented script (Python with BeautifulSoup, Scrapy or Selenium—use whichever fits best) accompanying the file would be appreciated (not mandatory) so I can rerun the extraction whenever the site updates.
Deliverables
• Excel (.xlsx) file containing the full data set
• Source script and brief usage notes
• Quick walkthrough of any environment setup needed to re-run the scrape
Once I verify the counts and spot-check a sample against the site, the job is done.
https://www2.gov.pt/fichas-de-enquadramento/fundacoes-e-pessoas-coletivas-de-utilidade-publica/pesquisa/-/pmc/lista
Accuracy is key: the sheet must mirror the live site, without missing or duplicated entries. If a page is paginated, please make sure all pages are traversed. Dynamic content that loads as you scroll should be handled as well.
A lightweight, well-commented script (Python with BeautifulSoup, Scrapy or Selenium—use whichever fits best) accompanying the file would be appreciated (not mandatory) so I can rerun the extraction whenever the site updates.
Deliverables
• Excel (.xlsx) file containing the full data set
• Source script and brief usage notes
• Quick walkthrough of any environment setup needed to re-run the scrape
Once I verify the counts and spot-check a sample against the site, the job is done.
https://www2.gov.pt/fichas-de-enquadramento/fundacoes-e-pessoas-coletivas-de-utilidade-publica/pesquisa/-/pmc/lista
Related categories:
JavaScript
Python
Excel
Web Scraping
Scrapy
Data Extraction
BeautifulSoup
Selenium