Tender Data Web Scraping

Job ID: 40618494

Budget: ₹1,500 – ₹12,500 INR

I have a set of URLs that point to various procurement and tender-posting sites, and I need every relevant tender detail extracted. You’ll be working exclusively from the links I supply, so there’s no site discovery involved—just clean, reliable scraping. One of the sites will require OTP and other sites will require captcha.

Requested data
The target pages list tender notices; for each notice I need the essential fields (title, reference number, issuing body, publication and closing dates, link to full notice, etc.) captured consistently across all sites. If additional obvious fields appear, include them as well.

Output format
Everything must arrive in a single, well-structured Excel spreadsheet, one row per tender with clear column headings. Please normalise dates and strip duplicates so the file is immediately usable for analysis.

Technical approach
Feel free to employ Python, Scrapy, BeautifulSoup, Selenium or any combination that handles pagination, dynamic content or CAPTCHA challenges. The key is repeatability: I’d like to rerun the script later, so provide clean, documented code.

Deliverables
• Final Excel workbook containing all scraped tender records
• Runnable script or notebook with brief setup instructions
• I will want the script to be running on daily basis and will also need it to be updated on any changes on the website in the future.
• The source code will belong to us and is not to be shared with anyone else

Accuracy, respectful site access, and timely delivery are paramount. When you reply, outline your proposed workflow, expected turnaround, and examples of similar scraping projects you’ve completed.