E-commerce Product Data Scraping Needed
Budget: ₹1,500 – ₹12,500 INR
I’m looking for a reliable web-scraping expert to pull product information from several e-commerce sites. The focus is on specific details and caselaws associated with each item, so accuracy and clean structuring matter.
Scope of work
• Build or configure a scraper that navigates target e-commerce pages and extracts the required fields.
• Capture the specific details and caselaws exactly as they appear on the site, along with any related product identifiers for easy cross-reference.
• Deliver the data in a tidy CSV or JSON file and include the reusable script (Python, Node, or your preferred language) plus brief run instructions.
• Respect site policies, handle pagination, and avoid IP blocks through polite request timing or proxy rotation.
Acceptance criteria
• All listed URLs are scraped without missing records.
• Each row contains: product name, key details, linked caselaw text or citation, page URL, date of capture, and any additional metadata you spot that improves traceability.
• No duplicate entries and all text is UTF-8 encoded.
A medium scope budget is set aside, so I’m expecting a solid, tested solution rather than a quick one-off scrape. If you’ve previously handled e-commerce data extraction or legal reference gathering, mention it and share a small sample or repo link. Looking forward to working together.
Scope of work
• Build or configure a scraper that navigates target e-commerce pages and extracts the required fields.
• Capture the specific details and caselaws exactly as they appear on the site, along with any related product identifiers for easy cross-reference.
• Deliver the data in a tidy CSV or JSON file and include the reusable script (Python, Node, or your preferred language) plus brief run instructions.
• Respect site policies, handle pagination, and avoid IP blocks through polite request timing or proxy rotation.
Acceptance criteria
• All listed URLs are scraped without missing records.
• Each row contains: product name, key details, linked caselaw text or citation, page URL, date of capture, and any additional metadata you spot that improves traceability.
• No duplicate entries and all text is UTF-8 encoded.
A medium scope budget is set aside, so I’m expecting a solid, tested solution rather than a quick one-off scrape. If you’ve previously handled e-commerce data extraction or legal reference gathering, mention it and share a small sample or repo link. Looking forward to working together.