News Website Text Scraping & Processing
Budget: €250 – €750 EUR
I'm looking for a skilled developer to scrape text contents from a specific news website. The project involves:
- Scraping text content from early 2022 to before 2024-11-22.
- Specifically extracting author names from the posts.
- Saving the source code of each post in separate folders named by the post’s date.
- Creating a second Python program to generate an Excel file based on the downloaded HTML.
You do not need to know how to extract Korean texts from images, as the texts are available. However, if you have experience in that area, feel free to include it in your bid.
Ideal skills and experience for the job include:
- Proficiency in Python, particularly for web scraping.
- Experience with data processing and Excel generation.
- Understanding of how to minimize server load while scraping.
Please note that the aim is to minimize the burden on the website, hence the suggestion of two separate Python programs. I'm open to any other efficient methods you may propose.
- Scraping text content from early 2022 to before 2024-11-22.
- Specifically extracting author names from the posts.
- Saving the source code of each post in separate folders named by the post’s date.
- Creating a second Python program to generate an Excel file based on the downloaded HTML.
You do not need to know how to extract Korean texts from images, as the texts are available. However, if you have experience in that area, feel free to include it in your bid.
Ideal skills and experience for the job include:
- Proficiency in Python, particularly for web scraping.
- Experience with data processing and Excel generation.
- Understanding of how to minimize server load while scraping.
Please note that the aim is to minimize the burden on the website, hence the suggestion of two separate Python programs. I'm open to any other efficient methods you may propose.