Scraping non-proprietary data

Job ID: 33697373

Budget: ₹37,500 – ₹75,000 INR

I am looking for programmers to scrape non-proprietary data (meta data and PDF documents) from government web sites. The salient requirements are –

1. Ability to take turn-key project: own data-scraping, not just write program.
2. New data keeps coming in, so ability to handle incremental data is must.
3. Initial proof of concept will be on local server – we will give you remote access to our server.
4. Final product to be delivered on AWS – meta data will be stored in Aurora and documents on S3.
5. Ability to parse PDF documents and apply proximity logic to identify meta data.
6. Data is huge – about 5 million records – so we need to run parallel services – probably 100+.
7. For running parallel services, data scraping must have beginning point and end point logic.
8. Ability to handle Captcha.
Related categories: Data Scraping