Enhanced word search in Google Scholar
Budget: $750 – $1,500 USD
Editors and scholar frequently use Google Scholar to check how frequently a given word or combination of words is used in the literature. Currently, this approach is not ideal for essentially two reasons: 1) there is an increasing number of articles from non-Native English speakers written in poor English; 2) Word/expression usage is often field-specific. Thus, I would like to have a program that performs two tasks:
1) The word/expression frequency would be searched only in articles that fulfill several conditions, namely: a) The surnames of the first and last authors are frequent in anglophone countries; b) The country of the institution of the first and last author is anglophone; 2) The journal is from a field that the person doing the search can indicate (e.g., Physics, Biology, Cardiology, Math, etc.); 3) The impact factor of the journal can be selected (e.g., above 3). Additional features along these lines may be required.
2) It would also be helpful to implement a systematic and automatic Google Scholar scan of the entire article using sliding windows of 1, 2, 3 or n consecutive words. The output (for a given expression wth n words) would be: a) a file with the entire text of the article highlighted with a color code for the frequency of each expression in Google Scholar; b) a list of the least frequent expressions (and their coordinates within the text). This would be extremely helpful to immediately identify the problematic passages. However, it is not a trivial task because for expressions above 3 or 4 words one would have to look for the frequency of the grammatical structure and not the exact combination of words (an exact number within the expression would have to become any given number and other similar adjustments would have to be made).
I would pay $1,000 for an online site with these features (if it works well) and a paywall. For the first two years after launching the page, the coder would also receive 40% of the profits from the fee the users pay to use the enhanced Google Scholar search. After 2 years, the coder would start receiving 10% of the profits.
1) The word/expression frequency would be searched only in articles that fulfill several conditions, namely: a) The surnames of the first and last authors are frequent in anglophone countries; b) The country of the institution of the first and last author is anglophone; 2) The journal is from a field that the person doing the search can indicate (e.g., Physics, Biology, Cardiology, Math, etc.); 3) The impact factor of the journal can be selected (e.g., above 3). Additional features along these lines may be required.
2) It would also be helpful to implement a systematic and automatic Google Scholar scan of the entire article using sliding windows of 1, 2, 3 or n consecutive words. The output (for a given expression wth n words) would be: a) a file with the entire text of the article highlighted with a color code for the frequency of each expression in Google Scholar; b) a list of the least frequent expressions (and their coordinates within the text). This would be extremely helpful to immediately identify the problematic passages. However, it is not a trivial task because for expressions above 3 or 4 words one would have to look for the frequency of the grammatical structure and not the exact combination of words (an exact number within the expression would have to become any given number and other similar adjustments would have to be made).
I would pay $1,000 for an online site with these features (if it works well) and a paywall. For the first two years after launching the page, the coder would also receive 40% of the profits from the fee the users pay to use the enhanced Google Scholar search. After 2 years, the coder would start receiving 10% of the profits.