Heritrix crawler + UI for warc

Job ID: 37396702

Budget: $10 – $100 USD

We look for Heritrix/pywb and warc expert to implement a archiving of a page
Goal is to have later offline browsable scraped content of websites.

Your job will be to implement the crawler to scrape based on a URLs.
It is required to be able to run the scraper daily, weekly, bi-weekly, monthly.
and the endusers shall be able to visit the scraped stated freely form their browser

Required to provide dockerized containers, which get fired up via docker-compose

Additional requirements:
It shall be used deduplication to reduce required storage (after the first run, the scraper shall check for changes on the page and persist the changed pages in new folders)
https://github.com/internetarchive/heritrix3/wiki/Deduping-%28Duplication-Reduction%29
https://pywb.readthedocs.io/en/latest/manual/recorder.html#deduplication-filters

Milestones
MS1:
Implement the docker compose with sample domains (we will share in chat) which get crawled
provide sample configurations for the crawler and provide a default ocnfiguration that crawler + UI will work after running docker compose

MS2:
implement deduplication

MS3 (optional):
provide REST api call from java springboot to add new domains to the crawler + UI

MS4 (optional):
implement screenshoting of the visited pages

Your background is:
- multiple years of experience with docker/docker compose
- multiple years of experience with web scraping

If you are a good fit, you are open to get more tasks about implementing solutions fully on your own (e.g. with your team)

Budget?
will not be disclosed, place your best bid to get considered

What is next?
We will share you a NDA and afterwards a paid test task.

Payment?
- you estimate in a WBS (optimistic, expected, pessimistic, where optimistic < expected < pessimistic) after getting the task
- we discuss about clearances and effort
- we mutually agree to effort
- we assign you the task after mutually agreed
- you implement & delivery
- we pay
(basically the rules of freelancer)

Closed book vs open book?
We work only on open book.
Closed book means you are unwilling to define a WBS for the work and you add only a price tag to the task.
We are sorry we will not hire you in such a case!

Deliveries?
- in our on premise git (access will be granted to you)
- full sources
- maven
- libs, need prior confirm and we prefer to use mostly latest stable versions
- JDK 17 (mostly LTS)
- CI/CD works fine with checkstyle, pmd, spotbugs
- works on our side too (stage + production)
Related categories: Web Scraping Docker Java Spring Docker Compose