Comprehensive Web Data Scraping for LLM

Job ID: 38599518

Budget: £20 – £250 GBP

I'm looking for an expert web scraper to extract all text content from a specific website, along with all the journal articles that the site references. The scraped data needs to be formatted in a way that's suitable for use in a Large Language Model (LLM).
The web site is in German and the results need to be in German and English

Key Requirements:
- Scrape all text content from a designated site.
- Identify and scrape all referenced journal articles.
- Format the scraped data suitably for an LLM.

Ideal Skills:
- Proficient in web scraping tools and techniques.
- Experienced in data formatting for machine learning purposes.
- Knowledgeable in handling and sourcing academic journal articles.
Related categories: Python Data Entry Web Scraping Web Search Data Mining