Remove Handwriting from Scanned Multilingual PDFs
Budget: $250 – $750 USD
Requirement Document
Project Title: Handwriting & Pen Mark Removal from Scanned PDF/Image (Preserve Printed Content)
Project Overview
I am looking for an experienced Computer Vision / OCR / AI developer who can develop a solution to automatically remove all handwritten content and pen-based markings from scanned PDF documents while preserving all printed content exactly as it is.
The solution should work on multilingual documents, especially:
• English
• Chinese (Simplified & Traditional)
The final output should be a clean PDF with all printed text, tables, images and layout preserved.
Functional Requirements
• Remove handwritten text, notes, comments, initials, signatures, numbers and paragraphs.
• Remove underlines (single/double/curved/zigzag).
• Remove highlights (yellow/green/blue/pink).
• Remove circles, rectangles, arrows, freehand drawings and scribbles.
• Remove check marks, crosses, strike-through and corrections.
• Remove margin notes, pen strokes, marker strokes and any pen/pencil annotations.
• Printed characters must NEVER be removed.
• Preserve printed English and Chinese text even if handwriting overlaps it.
• Preserve printed tables, borders, images, logos, QR codes and barcode.
• Support single and multi-page PDFs (100+ pages).
• Preserve page size, layout and resolution.
Edge Cases
• Blue handwriting over black text
• Black handwriting over black text
• English handwriting on Chinese document
• Chinese handwriting on English document
• Small handwriting
• Thick marker
• Overlapping annotations
1. Development preferably in Python.
2. 2 samples are in the Word document
3. Please provide past experience in doing OCR or similar handwriting projects or functions.
Project Title: Handwriting & Pen Mark Removal from Scanned PDF/Image (Preserve Printed Content)
Project Overview
I am looking for an experienced Computer Vision / OCR / AI developer who can develop a solution to automatically remove all handwritten content and pen-based markings from scanned PDF documents while preserving all printed content exactly as it is.
The solution should work on multilingual documents, especially:
• English
• Chinese (Simplified & Traditional)
The final output should be a clean PDF with all printed text, tables, images and layout preserved.
Functional Requirements
• Remove handwritten text, notes, comments, initials, signatures, numbers and paragraphs.
• Remove underlines (single/double/curved/zigzag).
• Remove highlights (yellow/green/blue/pink).
• Remove circles, rectangles, arrows, freehand drawings and scribbles.
• Remove check marks, crosses, strike-through and corrections.
• Remove margin notes, pen strokes, marker strokes and any pen/pencil annotations.
• Printed characters must NEVER be removed.
• Preserve printed English and Chinese text even if handwriting overlaps it.
• Preserve printed tables, borders, images, logos, QR codes and barcode.
• Support single and multi-page PDFs (100+ pages).
• Preserve page size, layout and resolution.
Edge Cases
• Blue handwriting over black text
• Black handwriting over black text
• English handwriting on Chinese document
• Chinese handwriting on English document
• Small handwriting
• Thick marker
• Overlapping annotations
1. Development preferably in Python.
2. 2 samples are in the Word document
3. Please provide past experience in doing OCR or similar handwriting projects or functions.