Data Scraper for Text and PDFs

Job ID: 38202197

Budget: €30 – €250 EUR

Website scrapper needed to extract some text and PDF files from one website.

GUI

There is a need for the ultra simple GUI, where I enter unique ID's in one column, and scrapper updates the columns next to it once data is extracted. One column is for the text data and other column is to note the progress (complete).

TARGET WEBSITE

Website has two layers of security scraper will have to pass:

a/ we have simple maths query; and
b/ accepting T&C's by clicking a tick box.

Once security is passed, unique ID is to be added to the search box, and results are to be extracted (Text + PDF files).

WHERE TO SAVE FILES

Once extract is completed, all PDF files are to be saved in VPS folder.

Folder name: Unique ID used for search.

Naming convention of the documents to be saved:

a_b_c_d

a = filing ID which will be found on the website next to PDF document
b = Unique ID
c = filing type*
d = filing details**

* = this is standardised text next to each PDF file found on the website. There is limited number of variations of the text and ONE unique value is applied.

** = this is standardised text next to each PDF file found on the website. There is limited number of variations of the text and MULTIPLE unique values can be applied.

* & ** = idea here would be to shorten the names of the files when saving them in the server and use unique ID's for c & d.

For example: instead of saving full name of the [c] filing type to the PDF that is extracted, we use unique value ID "1" for "Registration". "2" for "Modification".

When we scan the website, we if we see "Registration" then "1" would be added to the PDF document name when saving: a_b_"1"_d.

Where new value is found, system must create unique value for it. So first, scrapper checks the table of "c" entries, if match, then we use the number, otherwise create new entry and use unique ID for it for section "c" for the naming convention.

Same logic applies to "d", but here more than one value can be true. So there could be a case of a_b_c_1&2&4 where each number represents unique value found.

All these unique values are to be saved somewhere where I can read the data and pull the data from.

Instructions on how to access the website are attached.

HOW SCRAPPER WORKS

Scrapper should be hosted & run autonomously on VPS to which you will be granted access to.
Related categories: PHP Python Web Scraping Selenium Webdriver Selenium