Real Estate Data Extraction via Web Scraping

Job ID: 39600167

Budget: ₹12,500 – ₹37,500 INR

Scope of Work – Freelance Web Scraping
Project Title: Structured Web Scraping for Real Estate Listings
Deadline: 30 July
Timeline: awarded freelancer shall use best efforts to complete the project by the ideal target date of 15 July. The final deadline for full and satisfactory completion of the project is 30 July. Failure to complete the project in a fully functional state by the 30 July deadline shall constitute non-fulfillment of contractual obligations, resulting in the project being deemed unsuccessful and rated as zero (0) for acceptance purposes.

Objective
To develop a web scraper that extracts structured property data from a real estate platform. The process should begin from a search result page URL, from which the tool must scrape all listings under that search criteria.
The scraper must:
* Accurately map each listing to a predefined CSV format (see below).
* Download, rehost, and link all images.
* Enhance the Title and Description (Content) using AI for clarity and appeal.
* Use the source link as a unique identifier.
* Store data in a structured format to allow database-based export.
Important: The solution must be fully functional and validated using live production data. The project will be considered complete only after successful testing with real-world data and confirmation that all expected objectives are met in full. This is a binary acceptance: either it works completely as expected (1) or it does not (0). Anything less than the full working solution will not be accepted.

Deliverables
The final deliverable must be a fully functional and validated scraping tool that:
* Is hosted on our platform using our hosting services.
* Is deployed through our GitHub repository (access will be provided).
* Successfully scrapes and exports real-world data in the required structure.
* Meets all functional requirements outlined above without exceptions.
* Is tested against a set of predefined links from various domains (to be provided by us). Test completion is binary: either the tool performs as specified across all test cases (1), or it does not (0).

Notes
* All text fields must be cleaned: no broken tags, no inline JS, no malformed HTML.
* Set up smart delays, user-agent rotation, or proxy rotation if scraping blocks occur.
* Scraper must dynamically accept a search result page URL.
* The scraper should be robust enough to skip and log broken or missing listings.
Optional Notes
* A CSV template will be attached. Confidential search URLs will be provided later via chat for testing purposes.
* If successful, this project may lead to long-term developmental projects, including ongoing data scraping or system integration work.


Preferred Technical Skills:

Strong proficiency in web scraping frameworks and libraries (e.g., Python with Scrapy, BeautifulSoup, Selenium, Puppeteer).

Experience with handling dynamic websites and JavaScript-rendered content.

Knowledge of proxy management, user-agent rotation, and anti-blocking techniques.

Familiarity with image downloading, processing, and rehosting workflows.

Experience integrating AI-based text enhancement.

Ability to export data in structured CSV and database-friendly formats.

Experience with GitHub for version control and deployment workflows.

Familiarity with cloud hosting platforms and deployment (preferably the client’s hosting environment).

Strong debugging, logging, and error-handling practices to ensure robustness.



(Please find attached Original copy of this SOW and template)