Data Scraping Expert Needed
Budget: €30 – €250 EUR
I'm seeking a skilled data scraper to extract substantial economic specifics from the website https://www.fatturatoitalia.it/regione/ .
This directory contains the index of all the Italian companies, divided by Provinces.
There are 107 Provinces in Italy. Overall, total number of Italian companies is between 2,996,280 and 2,997,180.
Every Province index contains its companies list in a paginated way, 45 rows per page.
Every row contains a link to the company detail page, where key information are nested.
You need to scrape every Province, every paginated table for each Province index, every company detail for each row in each page of the Province index.
When nested information are not present, please put N/A.
Key Information to Extract:
- Company Name (from Province index)
- Region (from Province index)
- Province (from Province index)
- City (from Province index)
- Employees costs (from company detail page, when available)
- Revenues year 1 (from company detail page, when available)
- Profits year 1 (from company detail page, when available)
- Revenues year 2 (from company detail page, when available)
- Profits year 2 (from company detail page, when available)
- Revenues year 3 (from company detail page, when available)
- Profits year 3 (from company detail page, when available)
- Link to company detail page (from Province index)
Your ability to utilize various data scraping tools effectively to retrieve comprehensive and accurate data will be essential in this project. The successful candidate should possess strong tech skills including proficiency in Python, Beautiful Soup, or other relevant scraping tools. Prior experience associated with similar tasks will be valued. I am looking forward to accurate and well-structured data.
Please check the attached deck to appreciate the details needed to succeed with the project.
This directory contains the index of all the Italian companies, divided by Provinces.
There are 107 Provinces in Italy. Overall, total number of Italian companies is between 2,996,280 and 2,997,180.
Every Province index contains its companies list in a paginated way, 45 rows per page.
Every row contains a link to the company detail page, where key information are nested.
You need to scrape every Province, every paginated table for each Province index, every company detail for each row in each page of the Province index.
When nested information are not present, please put N/A.
Key Information to Extract:
- Company Name (from Province index)
- Region (from Province index)
- Province (from Province index)
- City (from Province index)
- Employees costs (from company detail page, when available)
- Revenues year 1 (from company detail page, when available)
- Profits year 1 (from company detail page, when available)
- Revenues year 2 (from company detail page, when available)
- Profits year 2 (from company detail page, when available)
- Revenues year 3 (from company detail page, when available)
- Profits year 3 (from company detail page, when available)
- Link to company detail page (from Province index)
Your ability to utilize various data scraping tools effectively to retrieve comprehensive and accurate data will be essential in this project. The successful candidate should possess strong tech skills including proficiency in Python, Beautiful Soup, or other relevant scraping tools. Prior experience associated with similar tasks will be valued. I am looking forward to accurate and well-structured data.
Please check the attached deck to appreciate the details needed to succeed with the project.