PDF Data Extraction

Job ID: 38516861

Budget: £250 – £750 GBP

I'm looking for a proficient developer to create a solution for extracting structured data from PDFs. It should support content in many languages, including Portuguese, French and English.

The application will need to identify and extract elements such as:
- Title
- Author
- Publication Date
- Images and Captions
- Articles and Links
- Descriptions
- References

The extracted data should be made available through APIs in a JSON format.

Ideal skills for this job include extensive experience with PDF data extraction, API development, and a strong command of JSON. Familiarity with magazine format PDFs will be an added advantage.