Python Journal Email Extraction

Job ID: 39783647

Budget: ₹600 – ₹1,500 INR

I have a set of journal-article URLs and need the publicly listed “corresponding author” email for each one. The full job is about 20,000 records, or roughly 1,000 per day if you prefer staged delivery. I’m flexible on the Python stack—BeautifulSoup, Scrapy, Selenium, or any combination—so long as it is robust and respects robots.txt.

For every address you capture, include the article title, DOI, and author name so I can trace it back later. Please deliver the results in Microsoft Word files arranged volume-wise: start a fresh document every 400 email records to keep each file a manageable size. A simple table in each Word file is fine; I’ll handle further formatting.

If you’re willing to share the scraper script plus a short read-me so I can rerun it, that’s a welcome bonus. Accuracy is critical, so only harvest email addresses that the journal site makes publicly available. A quick sample showing your approach will help speed up the decision.