Invoice Pipeline: Photo/PDF Pre-Processing → Nanonets (OCR with Review & Learning) → Webhook (CUG JSON)
Budget: €8 – €1,000 EUR
I need a developer to build a simple, robust pipeline that pulls invoices from one email inbox, applies automatic pre-processing to photos, sends documents to Nanonets (OCR with Review that learns from corrections), and—once approved—calls a webhook that outputs a clean, consistent CUG (line-level JSON) for my system.
Note: Formulas, supplier/customer matching, and business logic will be handled later in my core software. Here I only need a clear, precise CUG output.
Scope / Deliverables
Email intake (IMAP or similar): download attachments (PDF/JPG/PNG/HEIC).
Photo pre-processing (essential automatic pipeline) then submit PDFs/photos to Nanonets.
Nanonets: configure a line-items model with Review UI and learning from corrections; export via webhook for Approved content only.
Webhook (FastAPI or Node): receive Nanonets JSON → map to CUG (line-level), with basic normalization and simple dedup; deliver code + README (Docker welcome).
Handover session: how to make the model learn via Review, how to handle new suppliers, how to adjust the CUG mapping if fields change.
Collaboration (included in the quote)
During onboarding we’ll jointly agree practical solutions for the most common errors (kept high-level for now). The goal is to pass only reliable CUGs to the core. More specific topics (e.g., advanced multi-page handling) will be discussed later.
Requirements
Experience with Nanonets (or similar IDP), image pre-processing, and webhook/API development.
Clean code, clear README, and a short training at the end.
Budget
Under USD 1,000 total for the scope above.
How to Apply (brief)
1 link or short description of a similar project (IDP/OCR + review + webhook).
Your preferred stack for the webhook (FastAPI or Node).
Confirm the final training session and willingness to collaborate on common cases without surprise extras.
Code Delivery & Ownership (explicit)
Mandatory delivery of the full source code (pre-processing, Nanonets integration, CUG webhook) in a Git repository.
Include: Dockerfile/Compose, sample config, README with build/run steps, .env.example (no secrets), and simple deploy scripts.
Minimal tests: at least a webhook healthcheck/smoke test + sample files for pre-processing.
API documentation for the webhook (OpenAPI/Swagger preferred).
IP transfer: all code produced becomes my property; you confirm you can assign it and that it does not include components with incompatible licenses.
Credentials: share any keys/secrets securely; rotate/revoke at project end.
Handover/training: live session (screen-share) covering setup, Nanonets Review usage (so the model learns), onboarding new suppliers, and adjusting the CUG mapping.
Bug-fix window: reasonable small fixes after delivery are included for a short period (e.g., 14–30 days), without surprise extras.
Collaboration Style (Appendix)
Bug-fix and reasonable small corrections are included in scope, without surprise extras. For truly new functionality not covered by this scope, we’ll define micro-tasks with mini-quotes upfront. I’m looking for someone who will set a solid base that I’ll later extend with the software that consumes these CUGs.
Note: Formulas, supplier/customer matching, and business logic will be handled later in my core software. Here I only need a clear, precise CUG output.
Scope / Deliverables
Email intake (IMAP or similar): download attachments (PDF/JPG/PNG/HEIC).
Photo pre-processing (essential automatic pipeline) then submit PDFs/photos to Nanonets.
Nanonets: configure a line-items model with Review UI and learning from corrections; export via webhook for Approved content only.
Webhook (FastAPI or Node): receive Nanonets JSON → map to CUG (line-level), with basic normalization and simple dedup; deliver code + README (Docker welcome).
Handover session: how to make the model learn via Review, how to handle new suppliers, how to adjust the CUG mapping if fields change.
Collaboration (included in the quote)
During onboarding we’ll jointly agree practical solutions for the most common errors (kept high-level for now). The goal is to pass only reliable CUGs to the core. More specific topics (e.g., advanced multi-page handling) will be discussed later.
Requirements
Experience with Nanonets (or similar IDP), image pre-processing, and webhook/API development.
Clean code, clear README, and a short training at the end.
Budget
Under USD 1,000 total for the scope above.
How to Apply (brief)
1 link or short description of a similar project (IDP/OCR + review + webhook).
Your preferred stack for the webhook (FastAPI or Node).
Confirm the final training session and willingness to collaborate on common cases without surprise extras.
Code Delivery & Ownership (explicit)
Mandatory delivery of the full source code (pre-processing, Nanonets integration, CUG webhook) in a Git repository.
Include: Dockerfile/Compose, sample config, README with build/run steps, .env.example (no secrets), and simple deploy scripts.
Minimal tests: at least a webhook healthcheck/smoke test + sample files for pre-processing.
API documentation for the webhook (OpenAPI/Swagger preferred).
IP transfer: all code produced becomes my property; you confirm you can assign it and that it does not include components with incompatible licenses.
Credentials: share any keys/secrets securely; rotate/revoke at project end.
Handover/training: live session (screen-share) covering setup, Nanonets Review usage (so the model learns), onboarding new suppliers, and adjusting the CUG mapping.
Bug-fix window: reasonable small fixes after delivery are included for a short period (e.g., 14–30 days), without surprise extras.
Collaboration Style (Appendix)
Bug-fix and reasonable small corrections are included in scope, without surprise extras. For truly new functionality not covered by this scope, we’ll define micro-tasks with mini-quotes upfront. I’m looking for someone who will set a solid base that I’ll later extend with the software that consumes these CUGs.