Puppeteer - Website Crawling project

Job ID: 33772770

Budget: €250 – €750 EUR

Hi developers,
I hope you are all having a lovely day.

I need a crawler that crawls websites, most importantly the ads on these websites. The crawler screenshots the website, the ads and saves the target and source urls of the ads. The crawler is optimized for a small list of websites, it should be easy however if new website styles can be added easily. There exists a prototype that already implements these features, so this project will mostly be re-coding the prototype but better and more professional. The prototype is JavaScript, there can be many things copied from the prototype. The new crawler should be coded new and implemented in TypeScript. Also, the finished data is stored in AWS S3 and a database (Preferrably AWS DynamoDB).
There should be an API implemented (secured with Bearer Token only) that exposes the crawled data and lets the API consumer add new URLs to crawler.
It is very important, that the new crawler is reliable and the code should be as structured and modular as possible, to add logic, if e.g. a website is added. The crawler+API should run in Docker on an AWS EC2 instance.

I will provide you with the source code of the app, as soon as we start chatting.

Technologies required:
- TypeScript
- Puppeteer
- Node.js/Express
- AWS S3, EC2, DynamoDB
- Docker

Best regards,
Florian