Custom OCR Engine Development
Budget: $2 – $5 USD
I need an OCR component that plugs directly into our existing order-processing software and reliably reads supplier purchase orders that arrive as JPEG scans. The engine only has to recognise English, but accuracy must be high enough to extract line-item tables, dates, totals and supplier IDs without manual correction.
You are free to build on Tesseract, OpenCV, or a deep-learning stack such as TensorFlow or PyTorch, as long as the final module returns structured JSON we can map to our database.
Deliverables I expect:
• Source code with clear build instructions
• A small training/validation dataset plus instructions for expanding it
• API documentation and a quick demo script
• Brief report on achieved accuracy and how to retrain the model if vendor layouts change
I can supply a representative set of purchase-order images for development and testing the moment we start.
You are free to build on Tesseract, OpenCV, or a deep-learning stack such as TensorFlow or PyTorch, as long as the final module returns structured JSON we can map to our database.
Deliverables I expect:
• Source code with clear build instructions
• A small training/validation dataset plus instructions for expanding it
• API documentation and a quick demo script
• Brief report on achieved accuracy and how to retrain the model if vendor layouts change
I can supply a representative set of purchase-order images for development and testing the moment we start.