Tender Data Web Scraping
Budget: ₹1,500 – ₹12,500 INR
I have a set of URLs that point to various procurement and tender-posting sites, and I need every relevant tender detail extracted. You’ll be working exclusively from the links I supply, so there’s no site discovery involved—just clean, reliable scraping. One of the sites will require OTP and other sites will require captcha.
Requested data
The target pages list tender notices; for each notice I need the essential fields (title, reference number, issuing body, publication and closing dates, link to full notice, etc.) captured consistently across all sites. If additional obvious fields appear, include them as well.
Output format
Everything must arrive in a single, well-structured Excel spreadsheet, one row per tender with clear column headings. Please normalise dates and strip duplicates so the file is immediately usable for analysis.
Technical approach
Feel free to employ Python, Scrapy, BeautifulSoup, Selenium or any combination that handles pagination, dynamic content or CAPTCHA challenges. The key is repeatability: I’d like to rerun the script later, so provide clean, documented code.
Deliverables
• Final Excel workbook containing all scraped tender records
• Runnable script or notebook with brief setup instructions
• I will want the script to be running on daily basis and will also need it to be updated on any changes on the website in the future.
• The source code will belong to us and is not to be shared with anyone else
Accuracy, respectful site access, and timely delivery are paramount. When you reply, outline your proposed workflow, expected turnaround, and examples of similar scraping projects you’ve completed.
Requested data
The target pages list tender notices; for each notice I need the essential fields (title, reference number, issuing body, publication and closing dates, link to full notice, etc.) captured consistently across all sites. If additional obvious fields appear, include them as well.
Output format
Everything must arrive in a single, well-structured Excel spreadsheet, one row per tender with clear column headings. Please normalise dates and strip duplicates so the file is immediately usable for analysis.
Technical approach
Feel free to employ Python, Scrapy, BeautifulSoup, Selenium or any combination that handles pagination, dynamic content or CAPTCHA challenges. The key is repeatability: I’d like to rerun the script later, so provide clean, documented code.
Deliverables
• Final Excel workbook containing all scraped tender records
• Runnable script or notebook with brief setup instructions
• I will want the script to be running on daily basis and will also need it to be updated on any changes on the website in the future.
• The source code will belong to us and is not to be shared with anyone else
Accuracy, respectful site access, and timely delivery are paramount. When you reply, outline your proposed workflow, expected turnaround, and examples of similar scraping projects you’ve completed.
Related categories:
Python
Data Processing
Excel
Web Scraping
Data Mining
Scrapy
Data Extraction
BeautifulSoup
Selenium
Data Management