Heritrix crawler + UI for warc
Budget: $10 – $100 USD
We look for Heritrix/pywb and warc expert to implement a archiving of a page
Goal is to have later offline browsable scraped content of websites.
Your job will be to implement the crawler to scrape based on a URLs.
It is required to be able to run the scraper daily, weekly, bi-weekly, monthly.
and the endusers shall be able to visit the scraped stated freely form their browser
Required to provide dockerized containers, which get fired up via docker-compose
Additional requirements:
It shall be used deduplication to reduce required storage (after the first run, the scraper shall check for changes on the page and persist the changed pages in new folders)
https://github.com/internetarchive/heritrix3/wiki/Deduping-%28Duplication-Reduction%29
https://pywb.readthedocs.io/en/latest/manual/recorder.html#deduplication-filters
Milestones
MS1:
Implement the docker compose with sample domains (we will share in chat) which get crawled
provide sample configurations for the crawler and provide a default ocnfiguration that crawler + UI will work after running docker compose
MS2:
implement deduplication
MS3 (optional):
provide REST api call from java springboot to add new domains to the crawler + UI
MS4 (optional):
implement screenshoting of the visited pages
Your background is:
- multiple years of experience with docker/docker compose
- multiple years of experience with web scraping
If you are a good fit, you are open to get more tasks about implementing solutions fully on your own (e.g. with your team)
Budget?
will not be disclosed, place your best bid to get considered
What is next?
We will share you a NDA and afterwards a paid test task.
Payment?
- you estimate in a WBS (optimistic, expected, pessimistic, where optimistic < expected < pessimistic) after getting the task
- we discuss about clearances and effort
- we mutually agree to effort
- we assign you the task after mutually agreed
- you implement & delivery
- we pay
(basically the rules of freelancer)
Closed book vs open book?
We work only on open book.
Closed book means you are unwilling to define a WBS for the work and you add only a price tag to the task.
We are sorry we will not hire you in such a case!
Deliveries?
- in our on premise git (access will be granted to you)
- full sources
- maven
- libs, need prior confirm and we prefer to use mostly latest stable versions
- JDK 17 (mostly LTS)
- CI/CD works fine with checkstyle, pmd, spotbugs
- works on our side too (stage + production)
Goal is to have later offline browsable scraped content of websites.
Your job will be to implement the crawler to scrape based on a URLs.
It is required to be able to run the scraper daily, weekly, bi-weekly, monthly.
and the endusers shall be able to visit the scraped stated freely form their browser
Required to provide dockerized containers, which get fired up via docker-compose
Additional requirements:
It shall be used deduplication to reduce required storage (after the first run, the scraper shall check for changes on the page and persist the changed pages in new folders)
https://github.com/internetarchive/heritrix3/wiki/Deduping-%28Duplication-Reduction%29
https://pywb.readthedocs.io/en/latest/manual/recorder.html#deduplication-filters
Milestones
MS1:
Implement the docker compose with sample domains (we will share in chat) which get crawled
provide sample configurations for the crawler and provide a default ocnfiguration that crawler + UI will work after running docker compose
MS2:
implement deduplication
MS3 (optional):
provide REST api call from java springboot to add new domains to the crawler + UI
MS4 (optional):
implement screenshoting of the visited pages
Your background is:
- multiple years of experience with docker/docker compose
- multiple years of experience with web scraping
If you are a good fit, you are open to get more tasks about implementing solutions fully on your own (e.g. with your team)
Budget?
will not be disclosed, place your best bid to get considered
What is next?
We will share you a NDA and afterwards a paid test task.
Payment?
- you estimate in a WBS (optimistic, expected, pessimistic, where optimistic < expected < pessimistic) after getting the task
- we discuss about clearances and effort
- we mutually agree to effort
- we assign you the task after mutually agreed
- you implement & delivery
- we pay
(basically the rules of freelancer)
Closed book vs open book?
We work only on open book.
Closed book means you are unwilling to define a WBS for the work and you add only a price tag to the task.
We are sorry we will not hire you in such a case!
Deliveries?
- in our on premise git (access will be granted to you)
- full sources
- maven
- libs, need prior confirm and we prefer to use mostly latest stable versions
- JDK 17 (mostly LTS)
- CI/CD works fine with checkstyle, pmd, spotbugs
- works on our side too (stage + production)