Web Scraping Expert Needed for E-Commerce Product Data Extraction (24 Hour Turnaround)
Budget: ₹12,500 – ₹37,500 INR
We are seeking an experienced web scraping specialist for data extraction tool for an e-commerce marketplace. The goal is to extract structured product data from 1,000 product URLs while bypassing bot detection and restrictions.
This is a time-sensitive project with a required turnaround of 24-48 hours. If this is successful, this could lead to long-term high-frequency, large-scale data extraction.
Required Data Fields to Extract
The extracted data must match the e-commerce platform's native API response. The key fields include:
1️⃣ Item Basic Information
title → Product name
image → Primary product image
label_ids → Internal labels (e.g., featured, best seller)
flag → Product status flag
status → Product availability status
brand → Product brand
2️⃣ Item Price & Model Information
model_id → Unique ID for each variation
price_before_discount → Original price before discounts
price → Current selling price
price_stocks → Stock levels per price model
status → Model availability status
3️⃣ Item Delivery Information
shipping_fee_info → Shipping fees
ungrouped_channel_infos → List of available shipping channels
channel_promotion_infos → Discounts applied to shipping
4️⃣ Seller & Shop Information
vacation → Whether the shop is on vacation
last_active_time → Timestamp of last seller activity
account → username → Seller’s username
5️⃣ Coupon Information
shop_vouchers → Available vouchers
min_spend → Minimum spend required
used_price → Discounted price
percentage_used → % of voucher usage
start_time → Start time of the coupon
end_time → Expiry date
discount_value → Discount amount
discount_percentage → Discount percentage
discount_cap → Maximum discount limit
6️⃣ Promotion Information
bundle_deal_info & promotion_info → Active promotions
min_amount → Minimum purchase amount for promo
fix_price → Fixed promotional price
discount_percentage → Discount percentage
start_time → Promotion start date
end_time → Promotion end date
discount_value → Discount amount
wholesale → Wholesale pricing details
Key Responsibilities:
✔ Extract structured product data from a provided list of 1,000 product URLs.
✔ Implement anti-bot evasion techniques (proxy rotation, CAPTCHA bypass, session handling).
✔ Deliver data in JSON format, structured to match an existing API response format.
✔ Maintain at least an 95%+ success rate (i.e., minimal failed extractions).
✔ Ensure that the scraping solution can be scaled beyond this phase if needed.
✔ Provide error logs for any URLs that fail to return data.
Expected Deliverables:
? Extracted product data for 1,000 URLs in JSON format (must match a predefined API structure).
? A success rate report, highlighting successful and failed extractions.
? Basic logs and error handling details for troubleshooting.
? Delivery within 24-48 hours.
Technical Requirements:
Required Skills & Experience:
✅ Strong experience with web scraping frameworks (Scrapy, Selenium, Puppeteer, Playwright, BeautifulSoup, etc.).
✅ Expertise in bot detection evasion, including:
- Proxy rotation (residential or datacenter proxies).
- User-agent spoofing & session handling.
CAPTCHA solving (2Captcha, Anti-Captcha, or AI-based bypass).
✅ Ability to structure extracted data in JSON format, following a predefined schema.
✅ Experience scraping e-commerce websites (Amazon, Home Depot, Alibaba or similar).
✅ Familiarity with reverse engineering API responses (preferred but not required).
✅ Experience with cloud-based scraping (AWS Lambda, GCP, Azure) is a plus.
Bonus (Nice to Have):
⭐ Ability to build a scalable, high-volume scraping pipeline for future expansion.
⭐ Familiarity with asynchronous or batch-based data extraction methods (S3, API integration).
Project Constraints & Challenges:
⚠️ Bot Protection: The website uses anti-scraping measures (IP bans, rate limiting, bot detection). You must have experience overcoming these challenges.
⚠️ Speed & Stability: The script must complete 1,000 extractions with an 95%+ success rate.
⚠️ Fast Delivery: You must be able to complete this within 24-48 hours.
Next Steps:
If you have proven experience in web scraping & anti-bot evasion, please submit:
? A brief overview of your past scraping projects (especially e-commerce).
? Examples of structured JSON outputs from previous scraping jobs.
? A realistic timeline for completing this task within the required 24-48 hours.
Start your job post with [grass is green] for it to be considered.
Looking forward to your proposals!
This is a time-sensitive project with a required turnaround of 24-48 hours. If this is successful, this could lead to long-term high-frequency, large-scale data extraction.
Required Data Fields to Extract
The extracted data must match the e-commerce platform's native API response. The key fields include:
1️⃣ Item Basic Information
title → Product name
image → Primary product image
label_ids → Internal labels (e.g., featured, best seller)
flag → Product status flag
status → Product availability status
brand → Product brand
2️⃣ Item Price & Model Information
model_id → Unique ID for each variation
price_before_discount → Original price before discounts
price → Current selling price
price_stocks → Stock levels per price model
status → Model availability status
3️⃣ Item Delivery Information
shipping_fee_info → Shipping fees
ungrouped_channel_infos → List of available shipping channels
channel_promotion_infos → Discounts applied to shipping
4️⃣ Seller & Shop Information
vacation → Whether the shop is on vacation
last_active_time → Timestamp of last seller activity
account → username → Seller’s username
5️⃣ Coupon Information
shop_vouchers → Available vouchers
min_spend → Minimum spend required
used_price → Discounted price
percentage_used → % of voucher usage
start_time → Start time of the coupon
end_time → Expiry date
discount_value → Discount amount
discount_percentage → Discount percentage
discount_cap → Maximum discount limit
6️⃣ Promotion Information
bundle_deal_info & promotion_info → Active promotions
min_amount → Minimum purchase amount for promo
fix_price → Fixed promotional price
discount_percentage → Discount percentage
start_time → Promotion start date
end_time → Promotion end date
discount_value → Discount amount
wholesale → Wholesale pricing details
Key Responsibilities:
✔ Extract structured product data from a provided list of 1,000 product URLs.
✔ Implement anti-bot evasion techniques (proxy rotation, CAPTCHA bypass, session handling).
✔ Deliver data in JSON format, structured to match an existing API response format.
✔ Maintain at least an 95%+ success rate (i.e., minimal failed extractions).
✔ Ensure that the scraping solution can be scaled beyond this phase if needed.
✔ Provide error logs for any URLs that fail to return data.
Expected Deliverables:
? Extracted product data for 1,000 URLs in JSON format (must match a predefined API structure).
? A success rate report, highlighting successful and failed extractions.
? Basic logs and error handling details for troubleshooting.
? Delivery within 24-48 hours.
Technical Requirements:
Required Skills & Experience:
✅ Strong experience with web scraping frameworks (Scrapy, Selenium, Puppeteer, Playwright, BeautifulSoup, etc.).
✅ Expertise in bot detection evasion, including:
- Proxy rotation (residential or datacenter proxies).
- User-agent spoofing & session handling.
CAPTCHA solving (2Captcha, Anti-Captcha, or AI-based bypass).
✅ Ability to structure extracted data in JSON format, following a predefined schema.
✅ Experience scraping e-commerce websites (Amazon, Home Depot, Alibaba or similar).
✅ Familiarity with reverse engineering API responses (preferred but not required).
✅ Experience with cloud-based scraping (AWS Lambda, GCP, Azure) is a plus.
Bonus (Nice to Have):
⭐ Ability to build a scalable, high-volume scraping pipeline for future expansion.
⭐ Familiarity with asynchronous or batch-based data extraction methods (S3, API integration).
Project Constraints & Challenges:
⚠️ Bot Protection: The website uses anti-scraping measures (IP bans, rate limiting, bot detection). You must have experience overcoming these challenges.
⚠️ Speed & Stability: The script must complete 1,000 extractions with an 95%+ success rate.
⚠️ Fast Delivery: You must be able to complete this within 24-48 hours.
Next Steps:
If you have proven experience in web scraping & anti-bot evasion, please submit:
? A brief overview of your past scraping projects (especially e-commerce).
? Examples of structured JSON outputs from previous scraping jobs.
? A realistic timeline for completing this task within the required 24-48 hours.
Start your job post with [grass is green] for it to be considered.
Looking forward to your proposals!