Java scraping expert required
Budget: $10 – $150 USD
We look for scraping expert solving our requirements in Java.
Goal is to have later offline browsable scraped content of a section of websites.
Your job will be to implement the crawler to scrape based on a URL regex and to scrape the visited pages into a folder per page.
after the first run, the scraper shall check for changes on the page and persist the changed pages in new folders.
Typically each website-domain contains about 20 subpages, which are relevant to us.
Depending on the choosen path, you will require a specific solution like selenium, jsoup and so on.
Import is, that the content is offline browsable incl. the images!
Milestones
MS1:
Implement a simple crawler on a shared page which is downloading the sites for offline use.
It creates also screenshots of the visited pages, so that if something goes wrong the image view is at least available.
MS2:
- fix and enhance the scraper for more shared pages
MS3:
- implement a spring boot cron job, which runs your code on scheduled basis
Implementations:
- the impl interface for the next named service classes => share for approval
- service class(es), which solve the scraping, screenshoting, persisting, comparing
NO UI for now required!
NO database required!
No REST endpoint exposing for the above named methods required! (only consuming the apis in the given links)
Your background is:
- multiple years of experience with Java
- multiple years of experience with web scraping
- multiple years of experience with spring boot
If you are a good fit, you are open to get more tasks about implementing solutions fully on your own (e.g. with your team)
Budget?
will not be disclosed, place your best bid to get considered
What is next?
We will share you a NDA and afterwards a paid test task.
Payment?
- you estimate in a WBS (optimistic, expected, pessimistic, where optimistic < expected < pessimistic) after getting the task
- we discuss about clearances and effort
- we mutually agree to effort
- we assign you the task after mutually agreed
- you implement & delivery
- we pay
(basically the rules of freelancer)
Closed book vs open book?
We work only on open book.
Closed book means you are unwilling to define a WBS for the work and you add only a price tag to the task.
We are sorry we will not hire you in such a case!
Deliveries?
- in our on premise git (access will be granted to you)
- full sources
- maven
- libs, need prior confirm and we prefer to use mostly latest stable versions
- JDK 17 (mostly LTS)
- CI/CD works fine with checkstyle, pmd, spotbugs
- works on our side too (stage + production)
Goal is to have later offline browsable scraped content of a section of websites.
Your job will be to implement the crawler to scrape based on a URL regex and to scrape the visited pages into a folder per page.
after the first run, the scraper shall check for changes on the page and persist the changed pages in new folders.
Typically each website-domain contains about 20 subpages, which are relevant to us.
Depending on the choosen path, you will require a specific solution like selenium, jsoup and so on.
Import is, that the content is offline browsable incl. the images!
Milestones
MS1:
Implement a simple crawler on a shared page which is downloading the sites for offline use.
It creates also screenshots of the visited pages, so that if something goes wrong the image view is at least available.
MS2:
- fix and enhance the scraper for more shared pages
MS3:
- implement a spring boot cron job, which runs your code on scheduled basis
Implementations:
- the impl interface for the next named service classes => share for approval
- service class(es), which solve the scraping, screenshoting, persisting, comparing
NO UI for now required!
NO database required!
No REST endpoint exposing for the above named methods required! (only consuming the apis in the given links)
Your background is:
- multiple years of experience with Java
- multiple years of experience with web scraping
- multiple years of experience with spring boot
If you are a good fit, you are open to get more tasks about implementing solutions fully on your own (e.g. with your team)
Budget?
will not be disclosed, place your best bid to get considered
What is next?
We will share you a NDA and afterwards a paid test task.
Payment?
- you estimate in a WBS (optimistic, expected, pessimistic, where optimistic < expected < pessimistic) after getting the task
- we discuss about clearances and effort
- we mutually agree to effort
- we assign you the task after mutually agreed
- you implement & delivery
- we pay
(basically the rules of freelancer)
Closed book vs open book?
We work only on open book.
Closed book means you are unwilling to define a WBS for the work and you add only a price tag to the task.
We are sorry we will not hire you in such a case!
Deliveries?
- in our on premise git (access will be granted to you)
- full sources
- maven
- libs, need prior confirm and we prefer to use mostly latest stable versions
- JDK 17 (mostly LTS)
- CI/CD works fine with checkstyle, pmd, spotbugs
- works on our side too (stage + production)
Related categories:
Business, Accounting, Human Resources & Legal
Web Scraping
Java Spring
Selenium
Spring Boot