OCR for Text Extraction from Scanned PDFs

Job ID: 39100792

Budget: ₹37,500 – ₹75,000 INR

I'm looking for an expert in NLP and OCR to help with extracting text from my scanned documents.

1. Fax-to-Digital Conversion:
- Receive faxed medical imaging request forms via a fax API like e fax or Fax.Plus
- Convert faxes into digital formats (PDF or structured data) using OCR technology (Tesseract,
AWS Textract).
2. AI-Powered Data Processing:
- Extract and validate key data fields (patient name, imaging type, clinical indication and urgency) using NLP and AI.
- Standardize and store the processed data in a secure database.
3. Secure Dashboard:
- Provide a user-friendly interface for radiologist to:
- View and manage imaging requests.
- Filter and search requests by patient, imaging type, or status. - AI based protocols suggestion and editing.
- Ensure role-based access control (RBAC) for secure access.
4. Compliance and Security:
- Implement end-to-end encryption for data transmission and storage. - Maintain audit logs for all actions performed on the platform.
- Ensure compliance with PIPEDA and HIPAA regulations.
Ideal candidates for this project should have prior experience in working with OCR software, and have a solid understanding of NLP techniques for processing and analyzing the extracted text. Proficiency in handling PDF files is a must.
Related categories: OCR Node.js Artificial Intelligence NLP OpenAI