Advanced Document Processing Engineer Needed -- 2
Budget: $250 – $750 USD
Job Description
We are building a web application that allows users to upload documents and chat with their content.
The core challenge is accurate document content extraction at scale, with OCR used only when strictly necessary, and with precise bounding boxes to enable high-quality text highlights inside a PDF viewer.
This is not a basic OCR task.
The focus is precision, performance, low operational cost, and backend robustness.
We are looking for a senior-level engineer who understands document processing pipelines, OCR optimization, and production-ready backend systems.
________________________________________
Scope of Work (Milestone-Based)
•Milestone 0 – Technical Audit
Duration: 1–2 days
Deliverables:
•Review of current backend architecture
•Identification of technical and cost risks
•Proposed OCR + security architecture
•Clear, prioritized implementation plan
________________________________________
•Milestone 1 – Smart Document Extraction
Duration: 3–12 days
The system must handle PDFs and other document formats, including:
•.pdf, .doc, .docx, .ppt, .pptx, .odt, .odp, .txt, .rtf, .md, .html, .htm, .jpg, .jpeg, .png
Document & Page-Level Detection Strategy:
1.100% selectable text documents
•No OCR at all (zero Tesseract usage)
•Extract native text
•Generate bounding boxes from embedded text when possible
2.Mixed documents (text + scanned/image pages)
•OCR only the pages without selectable text
• Pages with text must never go through OCR
• Page-level state persistence
3. 100% scanned / image-based documents
• Avoid full Tesseract OCR for cost and performance reasons
• Use a low-cost vision AI to generate usable textual descriptions per page
• Output oriented to titles, sections, tables, and key fields
• Example: Gemini Flash / Flash-Lite or equivalent
________________________________________
• Milestone 2 – Text Highlights (Bounding Boxes)
Duration: 3–4 days
Implement precise highlights.
Key concepts:
• A bounding box is the exact rectangle enclosing a word or text fragment in page coordinates, not screen coordinates
• Highlights are visual overlays, not text selections
Workflow:
• Backend returns text and bounding boxes
• Frontend renders the PDF
• Frontend draws semi-transparent rectangles using bounding box coordinates
• Highlights must remain accurate with zoom and responsive layouts
________________________________________
• Milestone 3 – Critical Bugs (P0)
Duration: 3–5 days
• Authentication and login stability
• PDF upload flow
• Backend crashes
• Firestore security rules
• Storage rules
• App Check configuration
________________________________________
• Milestone 4 – Backend Protection
Duration: 2–3 days
• Rate limiting per user
• File size and page count validation
• Clear logging for debugging and monitoring
• Abuse prevention mechanisms
Without proper limits, a single user could upload thousands of documents and trigger massive OCR costs.
________________________________________
• Milestone 5 – Stability & Performance
Duration: 2–4 days
• Function optimization
• Reduced cold starts
• Improved error handling
• Overall backend reliability
________________________________________
Required Skills
• Strong backend experience (Node.js, Python, or similar)
• PDF processing and document parsing
• OCR systems (Tesseract or alternatives)
• Bounding boxes and coordinate systems
• Cost-aware cloud architecture
• Experience with scalable, production-grade systems
________________________________________
Nice to Have
• Google Cloud or Firebase experience
• Vision AI APIs (Gemini or similar)
• SaaS backend optimization
• Security and abuse-prevention strategies
We are building a web application that allows users to upload documents and chat with their content.
The core challenge is accurate document content extraction at scale, with OCR used only when strictly necessary, and with precise bounding boxes to enable high-quality text highlights inside a PDF viewer.
This is not a basic OCR task.
The focus is precision, performance, low operational cost, and backend robustness.
We are looking for a senior-level engineer who understands document processing pipelines, OCR optimization, and production-ready backend systems.
________________________________________
Scope of Work (Milestone-Based)
•Milestone 0 – Technical Audit
Duration: 1–2 days
Deliverables:
•Review of current backend architecture
•Identification of technical and cost risks
•Proposed OCR + security architecture
•Clear, prioritized implementation plan
________________________________________
•Milestone 1 – Smart Document Extraction
Duration: 3–12 days
The system must handle PDFs and other document formats, including:
•.pdf, .doc, .docx, .ppt, .pptx, .odt, .odp, .txt, .rtf, .md, .html, .htm, .jpg, .jpeg, .png
Document & Page-Level Detection Strategy:
1.100% selectable text documents
•No OCR at all (zero Tesseract usage)
•Extract native text
•Generate bounding boxes from embedded text when possible
2.Mixed documents (text + scanned/image pages)
•OCR only the pages without selectable text
• Pages with text must never go through OCR
• Page-level state persistence
3. 100% scanned / image-based documents
• Avoid full Tesseract OCR for cost and performance reasons
• Use a low-cost vision AI to generate usable textual descriptions per page
• Output oriented to titles, sections, tables, and key fields
• Example: Gemini Flash / Flash-Lite or equivalent
________________________________________
• Milestone 2 – Text Highlights (Bounding Boxes)
Duration: 3–4 days
Implement precise highlights.
Key concepts:
• A bounding box is the exact rectangle enclosing a word or text fragment in page coordinates, not screen coordinates
• Highlights are visual overlays, not text selections
Workflow:
• Backend returns text and bounding boxes
• Frontend renders the PDF
• Frontend draws semi-transparent rectangles using bounding box coordinates
• Highlights must remain accurate with zoom and responsive layouts
________________________________________
• Milestone 3 – Critical Bugs (P0)
Duration: 3–5 days
• Authentication and login stability
• PDF upload flow
• Backend crashes
• Firestore security rules
• Storage rules
• App Check configuration
________________________________________
• Milestone 4 – Backend Protection
Duration: 2–3 days
• Rate limiting per user
• File size and page count validation
• Clear logging for debugging and monitoring
• Abuse prevention mechanisms
Without proper limits, a single user could upload thousands of documents and trigger massive OCR costs.
________________________________________
• Milestone 5 – Stability & Performance
Duration: 2–4 days
• Function optimization
• Reduced cold starts
• Improved error handling
• Overall backend reliability
________________________________________
Required Skills
• Strong backend experience (Node.js, Python, or similar)
• PDF processing and document parsing
• OCR systems (Tesseract or alternatives)
• Bounding boxes and coordinate systems
• Cost-aware cloud architecture
• Experience with scalable, production-grade systems
________________________________________
Nice to Have
• Google Cloud or Firebase experience
• Vision AI APIs (Gemini or similar)
• SaaS backend optimization
• Security and abuse-prevention strategies