Internet Research & Scrapping (News data)

Job ID: 30675078

Budget: $10 – $30 USD

News related data scraping (English language only).

1. Generate a list of available/possible News categories and subcategories that various news feeds/sources have out there:

e.g. Nature, Disasters, Crime, Military, Politics, Religion, Health, Economics, Technology/Cyberspace (internet related), etc.

CSV file structure you should provide:
Category, sub-category

e.g.
"Business, Finance & Economics", "Accounting"
"Business, Finance & Economics", "Careers"
etc

Possible source/example: https://www.newsmaker.com.au/category/index
=> Find additional sources. Find as many categories and subcategories you can find. Use separate CSV files for different sources, keep track of sources.

2. Get a list of known News sources (usually RSS feeds) at least 200 world wide
e.g. CNN, ABCNews, NBCNews, Aljazeera, Forbes, TechRadar, Yahoo, NYTimes, WSJ, Fox News, Guardian, etc.

CSV file structure you should provide:
source, URL, RSS_URL

Many news agencies offer multiple feeds (politics, sports, etc.), those are not counted as individual sources. Focus on main RSS feeds that include mixed/all types of news.

Possible source/examples:
https://libguides.wlu.edu/c.php?g=357505&p=2412837
https://blog.feedspot.com/usa_news_rss_feeds/
=> Find additional sources. Double check feeds are up and running

Include "Data" + "Masta !@#" (without +"!@# chars) in your proposal [anti-spam/anti-bot]. You'll not be considered for this task without it. Please provide time estimate in your proposal.