Web Scraping of Financial Times news article archive, Save PDFs or text files

Job ID: 37757872

Budget: $30 – $250 USD

I need a skilled freelancer to create a tool or script that can efficiently scrape the Financial Times' archives. The main focus will be on extracting the full text of articles and saving them as PDF or text files. This tool is essential for my research, allowing me to analyze content over time.

Here is the search field for FT: https://www.ft.com/search?q=biodiversity

Note that this may require you creating an FT account (a trial subscription for 1 months is only 1 USD).

Here are the key words you should scrape for:

1.
biodiversity,
ecosystem,
conservation,
wildlife,
species,
habitat,
environment,
natural resources,
ecological diversity,
biodiversity loss,
endangered species,
biodiversity conservation

2.
sustainable innovation,
Green technology
Eco-friendly advancements
Sustainable development
Environmental innovation
Conservation innovation
Renewable solutions
Sustainable entrepreneurship
Clean energy breakthroughs
Circular economy initiatives
Climate-friendly inventions
Sustainable business practices
Low-carbon innovations
Resource-efficient solutions
Eco-innovations
Sustainable manufacturing advancements
Sustainable development
Environmental conservation
Renewable energy
Waste management
Clean technology
Green technology
Climate change mitigation
Energy efficiency
Carbon footprint reduction
Circular economy
Sustainable transportation
Eco-friendly materials
Water conservation
Sustainable agriculture
Green building
Low-impact manufacturing
Sustainable packaging
Ecological innovation
Emissions reduction
Sustainable product design

Please save the output for both searches in separate folders that you provide upon the completion of the project.

**Key Requirements:**

- **Experience with Web Scraping:** Proven track record of scraping protected or subscription-based websites without violating their terms of service.

- **Data Handling:** Capability to efficiently extract full article texts, ensuring the integrity and readability of the content in the PDF output.

- **PDF Generation:** Proficient in generating PDF files from scraped content, with a keen eye for preserving the original formatting as closely as possible.

- **Language and Libraries:** Familiarity with Python (preferred) or another suitable programming language for web scraping and PDF generation. Experience with libraries like BeautifulSoup, Scrapy for scraping, and ReportLab or similar for PDF creation is highly desirable.

**Ideal Skills and Experience:**

- Experience with the Financial Times or similar financial news websites.
- Strong understanding of HTML/CSS, possibly JavaScript if needed for scraping dynamic content.
- Experience creating scripts/tools that can navigate through archives based on dates or keywords.
- Knowledge in handling rate limits and using proxies if required to ensure compliance with website policies.
- Ability to work with me to define the scope, such as specific sections or time frames from the archive.

I look for reliability, communication, and the ability to deliver the tool/software within a reasonable timeframe. If you have a portfolio showcasing similar projects, please share it in your bid.

Your expertise will not only contribute to valuable research but also ensure the preservation of journalistic content in an accessible format.
Related categories: Web Scraping Data Scraping