Advanced Python Web Crawler for Content Scraping
Budget: $30 – $250 USD
I need a Python web crawler that scrapes content from websites, including text on the pages and following all links. The crawler should be able to handle dynamic loading of content (e.g. through JavaScript) and not be blocked by no follow tags.
Key Requirements:
- Scrape and compose all texts on any website
- Follow all links to create a map of links and titles to content
- Handle dynamically loaded content, scraping it as well
- Deliver the scraped data in a plain JSON format, so that it is easy to read and process
- The output should contain the text following all links, visible and usable by me
Skills and Experience:
- Proficiency in Python and experience with web scraping using Python libraries (like BeautifulSoup, Scrapy, etc.)
- Familiarity with handling dynamic content and no follow tags within web scraping
- Understanding of HTML and JavaScript
- Ability to create a JSON response that is well-structured and easy to understand.
Key Requirements:
- Scrape and compose all texts on any website
- Follow all links to create a map of links and titles to content
- Handle dynamically loaded content, scraping it as well
- Deliver the scraped data in a plain JSON format, so that it is easy to read and process
- The output should contain the text following all links, visible and usable by me
Skills and Experience:
- Proficiency in Python and experience with web scraping using Python libraries (like BeautifulSoup, Scrapy, etc.)
- Familiarity with handling dynamic content and no follow tags within web scraping
- Understanding of HTML and JavaScript
- Ability to create a JSON response that is well-structured and easy to understand.