Python Document Data Extraction Automation
Budget: ₹12,500 – ₹37,500 INR
The core of this job is Extraction: I need a reliable Python solution that pulls structured information from documents—think invoices, forms, multi-page PDFs—by calling the Claude Vision API.
Here is what I’m after:
• A clean, well-commented Python script (or small module) that submits documents to the Claude Vision endpoint, captures the response, and writes the extracted fields to CSV or a database table.
• Configuration for common document layouts so I can add new field mappings without touching the code.
• Basic error handling and logging so failed pages or API timeouts are easy to diagnose.
Acceptance will be simple: I will run the script on a folder of sample files; if the key fields land in the output exactly as they appear in the documents, the job is done. Feel free to lean on pandas, requests, pydantic—or any other lightweight libraries that speed things up—as long as setup stays under a standard requirements.txt.
If this sounds straightforward to you, let’s get started.
Here is what I’m after:
• A clean, well-commented Python script (or small module) that submits documents to the Claude Vision endpoint, captures the response, and writes the extracted fields to CSV or a database table.
• Configuration for common document layouts so I can add new field mappings without touching the code.
• Basic error handling and logging so failed pages or API timeouts are easy to diagnose.
Acceptance will be simple: I will run the script on a folder of sample files; if the key fields land in the output exactly as they appear in the documents, the job is done. Feel free to lean on pandas, requests, pydantic—or any other lightweight libraries that speed things up—as long as setup stays under a standard requirements.txt.
If this sounds straightforward to you, let’s get started.
Related categories:
PHP
Python
Software Architecture
MySQL
Data Extraction
API
API Development
Claude (Anthropic)
Claude Code