Website scraper in Python
Budget: €30 – €250 EUR
I am looking for a skilled developer to create a website scraper with a user interface. I am fine if it is an already working website scraper from another project as I can imagine many people have already done a similar solution.
Project requirements:
- There should be a config file with config parameter for initial website address to scrape.
- The scraper should follow links (except header and footer)
- There should be a config file with config parameter for the limit of how many links to scrape (example: max 100 pages)
- There should be a config file with config parameter for the depth of scraping from the original URL (Example: 3 levels deep). For example, if we scrape example.com/cars with a depth level of 1, we also scrape example.com/cars/electric but not example.com/cars/electric/blue
Delivery Format:
- The scraped data should be delivered in TXT format into AWS S3
Ideal Skills and Experience:
- Proficiency in web scraping techniques and tools.
- Strong knowledge of Python, HTML, CSS, and JavaScript.
- Proficiency with AWS Cloud Infrastructure
- The scraped websites should be saved as plain text .TXT files in S3
Detailed video: https://www.loom.com/share/01663e90a7664dc891dfd94bb5c94b11?sid=c2b3e601-995f-47d4-8cd9-bc0d79707a28
In your application please specify that you have watched the video
Project requirements:
- There should be a config file with config parameter for initial website address to scrape.
- The scraper should follow links (except header and footer)
- There should be a config file with config parameter for the limit of how many links to scrape (example: max 100 pages)
- There should be a config file with config parameter for the depth of scraping from the original URL (Example: 3 levels deep). For example, if we scrape example.com/cars with a depth level of 1, we also scrape example.com/cars/electric but not example.com/cars/electric/blue
Delivery Format:
- The scraped data should be delivered in TXT format into AWS S3
Ideal Skills and Experience:
- Proficiency in web scraping techniques and tools.
- Strong knowledge of Python, HTML, CSS, and JavaScript.
- Proficiency with AWS Cloud Infrastructure
- The scraped websites should be saved as plain text .TXT files in S3
Detailed video: https://www.loom.com/share/01663e90a7664dc891dfd94bb5c94b11?sid=c2b3e601-995f-47d4-8cd9-bc0d79707a28
In your application please specify that you have watched the video