Content Aggregation Scraping
Budget: $30 – $250 AUD
I'm looking for a skilled web scraper to download specific (publicly available) content from an Australian government website: acnc.gov.au.
The ACNC site is a repository of publicly accessible information about not-for-profit organisations. However, it's a dynamic site so cannot use ChatGPT to access it.
In the section "Search for a Charity", upon searching for a charity name you get a link "search results" which then allows you to access "Financials & Documents". In the "financials & documents" there will be posted the "Financial Report 2024".
I will provide an Excel file with a column of text fields that is the exact name of each organisation. I would like the freelancer to download the financial report for each organisation, and store all the reports in a shared drive from which I can download the reports. The Financial reports are normally pdfs of between 0.5MB and 5MB in size (varies). In that shared drive, each report should be labelled with the name of the organisation with the suffix "FY24" added.
There is a total of 161 unique names (organisations) in the list. Making a wild guess, this will be about 200-300MB of data in all. The names in the list should correspond exactly to the names used in the ACNC site but there might be minor differences around abbreviations in the organisation name (e.g. Inc. or Incorporated, or Ltd vs Limited)
Below are example screenshots from the ACNC website indicating home to navigate to the desired data
Requirements:
- Scrape a government website continuing public information
- Update frequency: this is a one-time exercise
Ideal Skills and Experience:
- Proficiency in web scraping tools and techniques
- Experience with data storage
- Ability to handle CAPTCHAs and anti-scraping measures
- Knowledge of content categorization and organization
Please provide estimated timeline. What shared drive do you propose to use?
The ACNC site is a repository of publicly accessible information about not-for-profit organisations. However, it's a dynamic site so cannot use ChatGPT to access it.
In the section "Search for a Charity", upon searching for a charity name you get a link "search results" which then allows you to access "Financials & Documents". In the "financials & documents" there will be posted the "Financial Report 2024".
I will provide an Excel file with a column of text fields that is the exact name of each organisation. I would like the freelancer to download the financial report for each organisation, and store all the reports in a shared drive from which I can download the reports. The Financial reports are normally pdfs of between 0.5MB and 5MB in size (varies). In that shared drive, each report should be labelled with the name of the organisation with the suffix "FY24" added.
There is a total of 161 unique names (organisations) in the list. Making a wild guess, this will be about 200-300MB of data in all. The names in the list should correspond exactly to the names used in the ACNC site but there might be minor differences around abbreviations in the organisation name (e.g. Inc. or Incorporated, or Ltd vs Limited)
Below are example screenshots from the ACNC website indicating home to navigate to the desired data
Requirements:
- Scrape a government website continuing public information
- Update frequency: this is a one-time exercise
Ideal Skills and Experience:
- Proficiency in web scraping tools and techniques
- Experience with data storage
- Ability to handle CAPTCHAs and anti-scraping measures
- Knowledge of content categorization and organization
Please provide estimated timeline. What shared drive do you propose to use?