OCR Engine & Document Management System (DMS)

Job ID: 39564224

Budget: $30 – $250 USD

1. Project Overview
This project involves building the OCR Engine and Document Management System
(DMS) to digitize approximately 13 million physical pages comprising newsletters,
books, and articles. The primary objectives include:
 High-resolution scanning
 OCR and ICR text extraction
 Metadata indexing
 Implementation of a robust Document Management System (DMS)
 Support for multilingual content in Albanian and Arabic
2. Input & Output Specifications
2.1 Input Materials
 Physical documents such as Newspapers, Newsletters, Books and Articles will
be scanned and given at 400–600 DPI resolution to process.
 Output of scanned files should be in:
o PDF Format
o Archived format (e.g., PDF/A or TIFF for long-term preservation)
3. OCR & ICR Capabilities
The proposed application will have a OCR & ICR Engine to extract the content from the
source documents.
3.1 OCR (Optical Character Recognition)
 Convert scanned images into editable digital text
 Languages supported: Arabic and Albanian
 OCR Accuracy : as per Industry standard (https://research.aimultiple.com/ocraccuracy/)
o Printed text: Accuracy range >95% accuracy.
o Printed media: Accuracy range: ~60% to ~90%
o Handwriting: Accuracy range: ~20% to ~96%.
 The original layout of the document must be retained, with editability preserved.
3.2 ICR (Intelligent Character Recognition)
 Capable of converting handwritten documents to editable text
 Handwriting in Arabic and Albanian to be supported
 Non-converted handwritten content must be available in editable format post
extraction
4. Document Management System (DMS) Requirements
4.1 Metadata Management
 System must support automatic extraction and indexing of 15–20 predefined
metadata fields
 Automatically populate these fields into the DMS post processing
 Metadata should be editable post-ingestion
4.2 Search & Retrieval Functionality
 Full-text search capability across all stored documents
 Keyword search should:
o Highlight all keyword matches within the original documents
o Show both text and associated image content in the search results
 Unmatched or unconverted data must also available for manual correction
4.3 Auto-Segmentation and Layout Analysis
 Auto-segment scanned documents based on:
o Main headings
o Sub-headings
o Columns
o Article bodies
o Layout analysis engine must automatically crop and align content
 Retain article layout post-OCR for editorial flexibility
4.4 Auto-Linking and Indexing
 Articles must be:
o Auto-linked with related content
o Indexed with metadata and tags for fast retrieval
 Cross-referencing of topics and stories within the DMS should be enabled
4.5 Summary Generation
 Generate and store a short summary or synopsis for each digitized document
 Store this summary in a designated metadata field within the DMS
5. Processing Capabilities
5.1 Batch Processing
 Enable bulk ingestion and processing of scanned files
 Allow for automated field extraction and content storage into structured folders
 Workflow must support error reporting
5.2 Exception Handling
 Identify and flag:
o Unrecognized characters
o Skipped or misaligned segments
o Metadata mismatches
 Highlight these for manual review and correction
6. Storage & Integration
 All OCR and ICR processed outputs must be stored in the centralized DMS
 System should be scalable to manage over 13 million records eƯiciently
 Archived documents must be retrievable in compliance with long-term storage
standards
Related categories: Python OCR