Web crawling/scraping

Job ID: 33166794

Budget: $30 – $250 AUD

Project name: Real Estate Data Crawling
Purpose: Collect real estate news on a specific website

Test Url: https://batdongsan.com.vn
Real Estate Categories need to be crawled:
• https://batdongsan.com.vn/ban-dat
• https://batdongsan.com.vn/ban-nha-rieng
• https://batdongsan.com.vn/ban-nha-mat-pho
• https://batdongsan.com.vn/ban-dat-nen-du-an

Basic information need to get from website:
• Post ID
• Post Date
• Total Sqm
• Sale Price
• Price/m²
• Google GPS (Latitude, Longitude)
• Url

Crawled data needs to be inserted to table “Products” in MySQL with columns:
• PropertyCode
• EntryDate
• TotalSqm
• Price_List1
• PriceSqm
• GPS
• Url

Programming language & Sourecode:
• PHP & MYSQL

Functions need:
- Allow to crawl data from a "Start d/m/Y" to "End d/m/Y"
- Skip data without Google GPS information
- Insert automatically into MySQL table with avoid duplicated records filtered by Post ID

Example:
• Post ID = 32671467
• Post Date = 10/03/2022
• Total Sqm (m²) = 17000
• Sale Price (million VND) = 1.3 * 17000 = 22,100
• Price (million VND)/m² = 1.3
• Google GPS (Latitude, Longitude) = 20.8224818,105.4986538
• Url = https://batdongsan.com.vn/ban-dat-xa-cao-son_1/chinh-chu-em-can-chuyen-nhuong-lo-voi-dien-tich-la-1-7ha-tai-luong-son-hoa-binh-pr32671467