Web Scraping 9gag for Memes

Job ID: 37785600

Budget: £2 – £5 GBP

An efficient and effective web scraper is required for extracting English text and images from memes on 9gag.com.

The successful freelancer will need to:

Step 1:
- Webscrape 1000 image memes with text on it. The more text the better.
- The text should be in English.
- Memes should originate from the users in Great Britain in 2023.
- All memes should be unique, not versions of one and the same.
- Exclude those containing censored words like "bi*ch", "f*ck", etc. in the text on memes.
- Create a methodology for downloading the corresponding images
- Thoroughly test the tool to ensure it's error-free

I need to store the database for later usage, so ideally organize them with data into excel as well.

Step 2:
- Precisely target and extract English text from the image memes
- The text should not be a mix of miscellaneous words, but in the forms of sentences as it appears in the memes, or separate lines, so that the context in understandable.
- Form a database which can be converted into a single corpus and the one which indicates from which meme is this text, for further checking and analysis

The sample is below


Ideal skill set for this job would include:

- Excellent experience in web scraping, specifically text and images
- Strong knowledge in handling and optimizing high volumes of data
- Solid expertise in data extraction from multimedia
- Strong testing and troubleshooting skills
- Ability to meet tight deadlines.