Italian Road Hauliers Data Scraper

Job ID: 39311605

Budget: €30 – €250 EUR

Web Scraping Project: Extract Data from Italian Road Hauliers Register
Project Overview
I need a skilled web scraper to extract comprehensive data from the Italian National Register of Road Hauliers (Albo degli Autotrasportatori) website: https://www.alboautotrasporto.it/
Specific Requirements
Data Source
Main site: https://www.alboautotrasporto.it/
Search form page: https://www.alboautotrasporto.it/web/portale-albo/imprese-iscritte
Required Data Extraction
First Level Data: Extract the complete list of companies for each province-municipality combination from the search form (under "SEDE" section)


The form has dropdown selectors for "province" and "comuni"
Each search will return a list of companies (potentially paginated)
Second Level Data: For each company in the list, extract detailed information by accessing their individual profile pages


These details include registration information, company name, VAT number, address, vehicle information, etc. (as shown in the second screenshot I can share with you)
Technical Requirements
Develop an automated scraping solution that can handle the entire process without manual intervention
The solution must be able to:
Navigate through all province-municipality combinations systematically
Handle pagination in the results pages
Access and extract detailed information for each company
Store data in a structured format (CSV, Excel, or database)
Handle any website security measures (rate limiting, CAPTCHAs, etc.)
Deliverables
Complete dataset of all companies and their details in an agreed format
Source code for the scraping solution with documentation
Instructions on how to run the scraper if needed in the future
Brief report on the data collection process, including any challenges encountered
Important Notes
The scraper must be efficient and should minimize the load on the target website
The solution should be robust against potential website structure changes
Please ensure compliance with the website's terms of service and relevant data protection regulations
Skills Required
Web scraping (BeautifulSoup, Scrapy, Selenium, or similar)
Python (preferred) or other suitable programming language
Experience with handling forms and dynamic content
Data cleaning and structuring
Ability to work with anti-scraping measures
Timeline
Please provide your estimated timeline for completing this project.
Budget
Please provide your bid based on the scope of work described above.
Note: I will share screenshots of the website interface and sample data structure with the selected freelancer.