Multilingual PDF Text Extraction

Job ID: 39754964

Budget: $250 – $750 USD

I need a meticulous freelancer to carry out تفريغ النصوص from several PDF files that mix Arabic, English, and occasionally a third language. The assignment is purely text extraction—no layout recreation or design work—yet I still expect the final text to mirror the original order, paragraph structure, and punctuation.

Scope of work
• Run reliable OCR on each PDF and proof-read the result to eliminate recognition errors, especially with right-to-left sections and diacritics.
• Deliver each source file as a separate, clean .docx (UTF-8), ready for my editors to use.

Acceptance criteria
• 99 %+ character accuracy across all languages.
• Paragraphs, headings, lists, and page breaks retained where they exist in the source.
• No mixed-up character encoding or garbled RTL/LTR sequences.

If you are comfortable with multilingual OCR tools like ABBYY FineReader, Tesseract, Adobe Acrobat Pro, or similar solutions, let me know which one you plan to use and why. Please also share a brief sample of comparable work and tell me how long you would need to finish roughly 50 pages.