Data from https://www.gsmarena.com/ Product Images and Product Specification without Copyright Issue
Budget: £20 – £250 GBP
I need a fast, clean scrape of gsmarena.com that gives me two things for a list of phone models I will share the moment we start:
• every technical specification that appears on each model’s page, and
• the associated images, saved in the highest resolution available.
I will be working with Samsung, Apple and Xiaomi devices first, but the script should be flexible enough to add more brands later. All content must be captured in a way that lets me reuse it without infringing anyone’s copyright, so please focus on material that is clearly published for public reference on the site.
Deliverables I expect
1. A CSV or JSON file per brand (your choice) containing the full spec table for each model.
2. A neatly organised image folder, one sub-folder per model, filenames matching the phone name.
3. Well-commented Python, Node, or similar scraping code so I can re-run the job when new models appear.
Acceptance criteria
• No missing fields from the spec tables.
• Images download without watermarks and match the model in both name and variant.
• Script completes without manual intervention and respects polite crawl delays.
If this sounds straightforward to you and you can turn it around quickly, let me know how soon you can have an initial dump ready and which language or libraries you prefer to use (BeautifulSoup, Scrapy, Puppeteer, Selenium, etc.).
• every technical specification that appears on each model’s page, and
• the associated images, saved in the highest resolution available.
I will be working with Samsung, Apple and Xiaomi devices first, but the script should be flexible enough to add more brands later. All content must be captured in a way that lets me reuse it without infringing anyone’s copyright, so please focus on material that is clearly published for public reference on the site.
Deliverables I expect
1. A CSV or JSON file per brand (your choice) containing the full spec table for each model.
2. A neatly organised image folder, one sub-folder per model, filenames matching the phone name.
3. Well-commented Python, Node, or similar scraping code so I can re-run the job when new models appear.
Acceptance criteria
• No missing fields from the spec tables.
• Images download without watermarks and match the model in both name and variant.
• Script completes without manual intervention and respects polite crawl delays.
If this sounds straightforward to you and you can turn it around quickly, let me know how soon you can have an initial dump ready and which language or libraries you prefer to use (BeautifulSoup, Scrapy, Puppeteer, Selenium, etc.).
Related categories:
PHP
Python
Data Processing
Web Scraping
Data Mining
Node.js
Scrapy
Data Extraction
BeautifulSoup
Selenium