AI-Enhanced Web Content Scraper

Job ID: 39714136

Budget: $15 – $25 USD

I’m looking to have a robust web-crawler built that focuses purely on content scraping. The tool should navigate large sites without being blocked, extract the full HTML or clean text as required, and immediately pass that data through an AI routine for enrichment—think automatic language detection, summarisation, keyword tagging, or any other clever post-processing step you recommend.

Here’s what I need to walk away with:
• A documented crawler (Python, Scrapy/Selenium/BeautifulSoup or a comparable stack) that can be re-pointed to new domains via a simple config file.
• An integrated AI module (NLP library, GPT API, or similar) that enhances each record and outputs structured JSON/CSV ready for downstream use.
• Clear instructions on deploying the solution on my own server plus sample runs proving reliability on at least two target sites.

Please highlight your relevant experience building similar scrapers or AI pipelines; links to prior projects or repos are ideal. If you have creative ideas for additional enrichment steps, let me know—this project is meant to evolve.