Advanced Web Scraping & Data Extraction Routine
Budget: $750 – $1,500 USD
I am wanting a routine that can open a specified web page and save the page to a file location.
The inputs to the routine will be a URL specifying the web location to search and a path designating where to store the file. Additional inputs will be an ini file containing the proxy information needed as well as a “fingerprint” ini or repository that will allow the above routine to cycle through any proxies in a random fashion and any related fingerprints of browsers.
This project needs to be command line capable of getting its inputs, and of returning a set of results when it has completed the request. The results shall be able to determine an invalid URL, a missing file, a file with a password (which causes a failed download) and a success flag. I expect to run this routinely from my own code and expect to make millions of downloads throughout a year’s time.
A simple Cntl-S and the addition of a file path in front of the saved filename dialog would work well. However, if you hit a site more than a couple of times an hour, without the proxy and fingerprinting you will get noticed and blocked from many websites.
I have about 16 million links for the www.digikey website. I would like to provide you the links through an API that will run on my local equipment. The input will include the link and a path to store the link.
The output should be the equivalent of hitting Ctrl-S. and typing in the path to send the file. In addition to the above the task must get past the normal efforts on the website to keep Bots out. I need to get past most of the bot blocks and an api solution would be better than the ai.vision software I had used a year ago.
As a test of the process, I expect to:
Add proxies to the proxy.ini file, Create and add browser fingerprints to its fingerprint.ini file. Once loaded I will test the code by starting the program from a command line, providing the URL and file path to save it. I will expect the data acquired to be located in the correct location as well as a file called Reults.txt that let me know the success or failure.
Sample pages:
https://www.digikey.com/en/products/detail/texas-instruments/LM1971MX-NOPB/366746
https://www.digikey.com/en/products/detail/onsemi/FAN3852UC16X/9829106
Sample files are attached with the results for the two part numbers above.
I am expecting a complete solution as above for use on my local computer. It will need to include the executable (if applicable), all source code and applicable documentation on its use.
Special requirements include: - Proficiency in web scraping and data extraction. - Experience in handling and organizing large volumes of web data.
The inputs to the routine will be a URL specifying the web location to search and a path designating where to store the file. Additional inputs will be an ini file containing the proxy information needed as well as a “fingerprint” ini or repository that will allow the above routine to cycle through any proxies in a random fashion and any related fingerprints of browsers.
This project needs to be command line capable of getting its inputs, and of returning a set of results when it has completed the request. The results shall be able to determine an invalid URL, a missing file, a file with a password (which causes a failed download) and a success flag. I expect to run this routinely from my own code and expect to make millions of downloads throughout a year’s time.
A simple Cntl-S and the addition of a file path in front of the saved filename dialog would work well. However, if you hit a site more than a couple of times an hour, without the proxy and fingerprinting you will get noticed and blocked from many websites.
I have about 16 million links for the www.digikey website. I would like to provide you the links through an API that will run on my local equipment. The input will include the link and a path to store the link.
The output should be the equivalent of hitting Ctrl-S. and typing in the path to send the file. In addition to the above the task must get past the normal efforts on the website to keep Bots out. I need to get past most of the bot blocks and an api solution would be better than the ai.vision software I had used a year ago.
As a test of the process, I expect to:
Add proxies to the proxy.ini file, Create and add browser fingerprints to its fingerprint.ini file. Once loaded I will test the code by starting the program from a command line, providing the URL and file path to save it. I will expect the data acquired to be located in the correct location as well as a file called Reults.txt that let me know the success or failure.
Sample pages:
https://www.digikey.com/en/products/detail/texas-instruments/LM1971MX-NOPB/366746
https://www.digikey.com/en/products/detail/onsemi/FAN3852UC16X/9829106
Sample files are attached with the results for the two part numbers above.
I am expecting a complete solution as above for use on my local computer. It will need to include the executable (if applicable), all source code and applicable documentation on its use.
Special requirements include: - Proficiency in web scraping and data extraction. - Experience in handling and organizing large volumes of web data.