Python Website Scraper
Budget: $50 – $0 AUD
Python dev for a program of work to build a robust, production-grade website scraper on AWS using Selenium. A proof of concept has already been built on BeautifulSoup but we are looking to move this to Selenium for image-to-text rather than html-to-text scraping.
The website requires logon. The scraper will need to simulate a real end user browser with authentication, simulated browser, randomised timing, request throttling, round-robin user accounts - all to minimise the chance of being bloked. There are two levels of pages to be scraped - an index page which lists posts (and is paginated) and then the post page (which is also paginated). All scrape attempts and their responses need to be logged with assets (HTML, image files) stored in S3 for files and recorded in DynamoDB tables. Scraper needs to be monitored and managed using AWS administgration infrasutrcture (for logging, control, deployment, etc). Scraper will need to call into AWS OCR to parse message images to text. Scraper will be run on polling timer, minutely, within certain weekday hours.
We are looking for a developer that can then go further into parsing scraped content, either with simpler text fragment extraction (this has already been done in our early proof of concept), AWS machine learning or using mechanical turk for human parsing/verification.
The end result will be an AWS event or message queue reflecting the new data, stored in DynamoDB but placed into queue/event for consumers to action in our broader app architecture.
The website requires logon. The scraper will need to simulate a real end user browser with authentication, simulated browser, randomised timing, request throttling, round-robin user accounts - all to minimise the chance of being bloked. There are two levels of pages to be scraped - an index page which lists posts (and is paginated) and then the post page (which is also paginated). All scrape attempts and their responses need to be logged with assets (HTML, image files) stored in S3 for files and recorded in DynamoDB tables. Scraper needs to be monitored and managed using AWS administgration infrasutrcture (for logging, control, deployment, etc). Scraper will need to call into AWS OCR to parse message images to text. Scraper will be run on polling timer, minutely, within certain weekday hours.
We are looking for a developer that can then go further into parsing scraped content, either with simpler text fragment extraction (this has already been done in our early proof of concept), AWS machine learning or using mechanical turk for human parsing/verification.
The end result will be an AWS event or message queue reflecting the new data, stored in DynamoDB but placed into queue/event for consumers to action in our broader app architecture.