PDF Data Extraction & Validation Automation

Job ID: 39038404

Budget: $30 – $250 CAD

Exercise: PDF Processing Workflow for Field Extraction and Validation
Scenario:

Imagine Deeded wants to automate part of its document processing workflow. A PDF document containing client information (e.g., name, address, and mortgage details) needs to be uploaded. Certain fields must be extracted, validated by a human, and then have the key data fields stored in a Google Sheet for future use.

Objective:

Design a solution for this process and implement a basic prototype or provide a detailed explanation of the steps you would take. Use a no-code automation tool where possible (eg: n8n)

Instructions:
Doc Upload

Upload a PDF. You can take examples PDF from the internet.

Extract key pieces of data (e.g., Name, Address, Mortgage Amount, mortgage number, lender name).

Outline assumptions or edge cases you’d consider (e.g., handling missing fields, formatting issues).

Doc Validation

Create a simple interface where a human can validate the extracted fiends against the original document. A human should be able to edit a field before it is passed on to the google sheet.

For PDF documents that are dates more than 30 days ago, create a rejection workflow, whereas the customer receives a notification that their document is out of date and asking them to upload new version.
Related categories: Python API Google Cloud Platform Automation Zapier