Google Results Scraper with Proxies and Captchas Enabled
Budget: $30 – $250 USD
I am looking for a skilled developer to create a Google Results Scraper with Proxies and Captchas Enabled. This project requires expertise in web scraping and data extraction.
Specific search terms:
- The scraper should use search terms from a file called keywords.txt, with one keyword per line.
Preferred search engine:
- The scraper should be designed to scrape results from Google.
Scraped results format:
- The results should be provided in CSV format.
Ideal skills and experience for the job:
- Proficiency in web scraping and data extraction techniques.
- Experience working with proxies and captchas.
- Familiarity with file input and output operations.
- Knowledge of Google's search engine and its search result structure.
- Attention to detail to ensure accurate and reliable scraping results.
If you have the necessary skills and experience, please submit your proposal.
------------------Above is Freelancer's own Description of this Project-------
------------------Below is My Own Description of this Project------------------
Brief Description/Summary
I need a software application which can search Google using search operators to find out how much competition there is in a given set of keyword phrases (search terms).
The software will have the ability to use private proxies, given the relatively large number of searches which may result.
Detailed Description:
The software will run from a folder called ‘Keyword Competitiveness’.In the folder will be the software or whatever form it takes and an input file called keywords.txt.
There should be a provision for adding private proxies so that the search process is not interrupted too much. The software will read a file in the same folder called keywords.txt which should list one keyword phrase per line.
The software will copy the 1st keyword phrase in the input list, and paste it into google.com (but can you make it optional to search all the Google networks; i.e. google.co.uk, and all country TLDs).
The software will perform the following searches:
Allintitle:”KEYWORD PHRASE”
Allinurl:”KEYWORD PHRASE”
Allintitle:”KEYWORD PHRASE” allinurl:”KEYWORD PHRASE”
The software will also detect the Moz DA score of the 1st place organic result (I use the free Moz toolbar to do this).
The software will save the number of results returned to an output CSV file with five columns, which are:
| Keyword Phrase | Allintitle Number | Allinurl Number | Allintitle and Allinurl Number | 1st Place Moz Rank |
The software will then go to the next keyword phrase in the keywords.txt file and do exactly the same thing.
It will continue until the last keyword phrase in the keywords.txt file and then it will stop, and output the results CSV file.
There should also be a section where the frequency of searches is controlled so as not to incur stoppages or captchas, and to emulate human behaviour.
In this section there will be an input data field for Minimum Pause Between Searches, and another input data field for Maximum Pause Between Searches.
The user will enter the minimum and maximum pause time in seconds, and the software will choose a random time between these two, and pause for as long as this random number indicates. At the end of this pause it will resume the next search.
There will also be another input data field for the user to enter the maximum number of searches done before an option to stop for | 1 hour | 2 hours | … up to 24 hours |
(So, for example, if this is set to 24 hours, the software will do n searches, as input by the user, above, and then pause for 24 hours.)
Lastly, I think I had better cater for any captchas that appear. So there should be an option to use captcha solver resources (a switch for Yes/No). I have GSA Captcha Breaker. So the software needs to be able to use that to break captchas.
If the GSA Captcha Breaker cannot solve a captcha it should go to my Death By Captcha account for it to be solved by a human.
And if Death by Captcha cannot solve it then the software should simply stop and display the captcha and then I will try to solve it manually myself.
If I can’t solve it then the software will indicate on the output CSV file that it was not solved. (I can then look up the search numbers manually for any of these later on.)
That’s it!
Specific search terms:
- The scraper should use search terms from a file called keywords.txt, with one keyword per line.
Preferred search engine:
- The scraper should be designed to scrape results from Google.
Scraped results format:
- The results should be provided in CSV format.
Ideal skills and experience for the job:
- Proficiency in web scraping and data extraction techniques.
- Experience working with proxies and captchas.
- Familiarity with file input and output operations.
- Knowledge of Google's search engine and its search result structure.
- Attention to detail to ensure accurate and reliable scraping results.
If you have the necessary skills and experience, please submit your proposal.
------------------Above is Freelancer's own Description of this Project-------
------------------Below is My Own Description of this Project------------------
Brief Description/Summary
I need a software application which can search Google using search operators to find out how much competition there is in a given set of keyword phrases (search terms).
The software will have the ability to use private proxies, given the relatively large number of searches which may result.
Detailed Description:
The software will run from a folder called ‘Keyword Competitiveness’.In the folder will be the software or whatever form it takes and an input file called keywords.txt.
There should be a provision for adding private proxies so that the search process is not interrupted too much. The software will read a file in the same folder called keywords.txt which should list one keyword phrase per line.
The software will copy the 1st keyword phrase in the input list, and paste it into google.com (but can you make it optional to search all the Google networks; i.e. google.co.uk, and all country TLDs).
The software will perform the following searches:
Allintitle:”KEYWORD PHRASE”
Allinurl:”KEYWORD PHRASE”
Allintitle:”KEYWORD PHRASE” allinurl:”KEYWORD PHRASE”
The software will also detect the Moz DA score of the 1st place organic result (I use the free Moz toolbar to do this).
The software will save the number of results returned to an output CSV file with five columns, which are:
| Keyword Phrase | Allintitle Number | Allinurl Number | Allintitle and Allinurl Number | 1st Place Moz Rank |
The software will then go to the next keyword phrase in the keywords.txt file and do exactly the same thing.
It will continue until the last keyword phrase in the keywords.txt file and then it will stop, and output the results CSV file.
There should also be a section where the frequency of searches is controlled so as not to incur stoppages or captchas, and to emulate human behaviour.
In this section there will be an input data field for Minimum Pause Between Searches, and another input data field for Maximum Pause Between Searches.
The user will enter the minimum and maximum pause time in seconds, and the software will choose a random time between these two, and pause for as long as this random number indicates. At the end of this pause it will resume the next search.
There will also be another input data field for the user to enter the maximum number of searches done before an option to stop for | 1 hour | 2 hours | … up to 24 hours |
(So, for example, if this is set to 24 hours, the software will do n searches, as input by the user, above, and then pause for 24 hours.)
Lastly, I think I had better cater for any captchas that appear. So there should be an option to use captcha solver resources (a switch for Yes/No). I have GSA Captcha Breaker. So the software needs to be able to use that to break captchas.
If the GSA Captcha Breaker cannot solve a captcha it should go to my Death By Captcha account for it to be solved by a human.
And if Death by Captcha cannot solve it then the software should simply stop and display the captcha and then I will try to solve it manually myself.
If I can’t solve it then the software will indicate on the output CSV file that it was not solved. (I can then look up the search numbers manually for any of these later on.)
That’s it!