PubChem Data Automation
Budget: ₹2,500 – ₹3,000 INR
Technology:
Technology must be feasible for scraping big data, the said scraping tool should not be hanged, we are using with .net, or python Technology based some tools etc. You can also suggest some others based on the best fit for our case. Tool should not be hanged.
I need an expert in web scraping tools to automate a repetitive task for me. The website to scope is PubChem (https://pubchem.ncbi.nlm.nih.gov/).
Key Requirements: Please also refer attached excel
Data can search by keyword provided in excel.
1. Chemical Name - e.g. Sitagliptin
2. IUPAC Name - e.g. N,N-Diallyl-2-chloroacetamide
3. CAS RN. - e.g. 93-71-0
Note:
1. Only first best match, we need to collect, and out all searched compounds retrieve entries, we need
2. This file should be saved with GBRN+PubChem CID followed by .png or .jpg extension
3. When you collect the data in scraping excel sheet, Keyword, GBRN and TRRN, Ref. No. should also come with data so that we could identify the data basis of Keyword.
- Automate the extraction process with the final goal of a constant update of data.
- The data to be extracted includes specific details on chemical compounds information, their associated biological activities, and their molecular structures.
- Ensure the tool is able to handle large amounts of data without crashing.
The ideal freelancer for this job would have solid experience in web scraping, bot programming, and data extraction. Knowledge in chemistry or a related field would be a definite plus to accurately comprehend the nature of the data to be extracted. As this project necessitates precision and consistency, I require someone diligently detail-oriented.
Technology must be feasible for scraping big data, the said scraping tool should not be hanged, we are using with .net, or python Technology based some tools etc. You can also suggest some others based on the best fit for our case. Tool should not be hanged.
I need an expert in web scraping tools to automate a repetitive task for me. The website to scope is PubChem (https://pubchem.ncbi.nlm.nih.gov/).
Key Requirements: Please also refer attached excel
Data can search by keyword provided in excel.
1. Chemical Name - e.g. Sitagliptin
2. IUPAC Name - e.g. N,N-Diallyl-2-chloroacetamide
3. CAS RN. - e.g. 93-71-0
Note:
1. Only first best match, we need to collect, and out all searched compounds retrieve entries, we need
2. This file should be saved with GBRN+PubChem CID followed by .png or .jpg extension
3. When you collect the data in scraping excel sheet, Keyword, GBRN and TRRN, Ref. No. should also come with data so that we could identify the data basis of Keyword.
- Automate the extraction process with the final goal of a constant update of data.
- The data to be extracted includes specific details on chemical compounds information, their associated biological activities, and their molecular structures.
- Ensure the tool is able to handle large amounts of data without crashing.
The ideal freelancer for this job would have solid experience in web scraping, bot programming, and data extraction. Knowledge in chemistry or a related field would be a definite plus to accurately comprehend the nature of the data to be extracted. As this project necessitates precision and consistency, I require someone diligently detail-oriented.