Social Media and Blog Historical Data Web Scraping Project
Budget: £20 – £250 GBP
I'm looking for a skilled web scraper who can extract historical data (pre-2024) from specific Twitter, Reddit, and Facebook accounts and forums. The task also includes scraping popular posts containing specific keywords on these platforms, as well as data from certain independent blogs.
Key Requirements:
- historical (pre-2024) social media platform content & engagement by specific accounts, pages, groups, or keywords
Additional Information:
- All the specific accounts, forums, and keywords, have been identified.
- Summary of how data was scraped will be required for
- Need both raw and basic cleaned copies of data.
- Employer is an SME in topic that data is being collected about, and is available for consultation during scraping project.
- There is also a potential opportunity for a secondary project conducting similar work depending on results of initial data analysis.
Project Details:
Twitter Accounts (~300 accounts)
- Extracting pre-2024 historical data of specific accounts
- specific account profile details, including follower and following list
- accounts posts and timestamp + engagement statistics
Twitter Keywords (<15 keywords)
- extracting top 20% of posts with most engagement by specific keywords pre-2024
Reddit ( <20 subreddits)
- extracting top 20% of posts pre-2024 on a monthly basis from specific subreddits
Facebook (<10 pages & groups)
- Pre-2024 historical data of posts with the most engagement on an monthly/annual basis (activity-dependent) from specific groups & pages
Facebook Keywords (<15 keywords)
- extracting top 20% of posts pre-2024 with most engagement by specific keywords
Misc Blogs (<10)
- extraction of pre-2024 historical blog post content from specific blogs
Key Requirements:
- historical (pre-2024) social media platform content & engagement by specific accounts, pages, groups, or keywords
Additional Information:
- All the specific accounts, forums, and keywords, have been identified.
- Summary of how data was scraped will be required for
- Need both raw and basic cleaned copies of data.
- Employer is an SME in topic that data is being collected about, and is available for consultation during scraping project.
- There is also a potential opportunity for a secondary project conducting similar work depending on results of initial data analysis.
Project Details:
Twitter Accounts (~300 accounts)
- Extracting pre-2024 historical data of specific accounts
- specific account profile details, including follower and following list
- accounts posts and timestamp + engagement statistics
Twitter Keywords (<15 keywords)
- extracting top 20% of posts with most engagement by specific keywords pre-2024
Reddit ( <20 subreddits)
- extracting top 20% of posts pre-2024 on a monthly basis from specific subreddits
Facebook (<10 pages & groups)
- Pre-2024 historical data of posts with the most engagement on an monthly/annual basis (activity-dependent) from specific groups & pages
Facebook Keywords (<15 keywords)
- extracting top 20% of posts pre-2024 with most engagement by specific keywords
Misc Blogs (<10)
- extraction of pre-2024 historical blog post content from specific blogs